Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

The reliability of the serum dioxin measurement in veterans of Operation Ranch Hand.

A study was conducted on the reliability of the serum dioxin measurement of enlisted Ranch Hands veterans participating in the Air Force Health Study using paired serum dioxin measurements. The 46 veterans were not randomly selected, but their demographic characteristics, health, and dioxin levels were similar to those of 404 other enlisted Ranch Hand veterans who had a single dioxin measurement made in 1987. The average time between the measurements was 0.61 years, the first measurement made from blood drawn on 10 April 1987 and the second from blood collected at a subsequent physical examination. In original unit, the coefficient of reliability was 0.87 (95% confidence interval: 0.76, 0.94) when the first measurement was at or below 50 parts per trillion. The measurement had no reliability in original units when the first measurement was greater than 50 parts per trillion. After a logarithmic transformation, the coefficient of reliability was 0.96 (95% confidence interval: 0.93 to 0.98). These results suggest that the serum dioxin measurement should not be used in original units for any purpose when the value exceeds 50 parts per trillion. The measurement is, however, highly reliable after a logarithmic transformation over the entire range of concentrations. Other studies using the same analytical method to measure dioxin in serum could similarly benefit if the measurement used is on the natural logarithm scale.

Aerospace Medicine↗

Clinical measurement of testicular volume in adolescents: comparison of the reliability of 5 methods.

PURPOSE: Measurement of the testis is a more readily available method of estimating spermatogenesis. Doubt remains about the best instrument for measuring testicular volume. Lack of bias or accuracy of instruments has received too much emphasis in some studies, while to our knowledge no one has yet appropriately compared reliability statistically. We propose a simple new method for measuring testicular size based on visual comparison with graphic models, and describe the reliability and bias of this and 4 traditional methods. MATERIALS AND METHODS: Measurements of 42 adolescent testes were made in a certain sequence: graphic method, dimensional measurement, Prader orchidometer, ring orchidometer and ultrasound with ultrasound assumed to be the standard. Statistical analysis was based on the linear structural model. RESULTS: Statistical tests indicated that all 5 methods are equally reliable (R > 0.9). Although they are not equally accurate, actual testicular size can be calculated using each of these 5 methods and the equations of the linear structural model. CONCLUSIONS: The new graphic method proposed in this study is as reliable as other well-known methods for measuring testicular size. Actual testicular volume can be estimated without bias and with equal reliability from any of the 5 methods using the equations of the linear structural model. This statistical approach is more relevant than the sole comparison of lack of bias or accuracy, which has been the main concern of previous studies.

Adolescent↗

Levels of reliability in fertility survey data.

A number of factual and attitudinal questions asked in the 1973 Taiwan KAP-4 survey were repeated in a postenumeration survey one month later in order to assess the reliability of responses of the 286 women reinterviewed. The level of reliability is found to vary depending on the measures used and on whether the focus is aggregate data or individual responses. Analysis of consistency of responses shows that while overall reliability at both aggregate and individual levels is reasonably good, there is greater reliability for factual than for attitudinal data. Nevertheless, consistency of responses on factual questions varies considerably depending on the salience of the topic to the respondent. Estimates of reliability are shown to depend on the measure used and on the skewness of the distributions of the responses.

Contraception↗

Reliability of the gross motor function measure in cerebral palsy.

The Gross Motor Function Measure (GMFM), an instrument comprising five dimensions devised by Russell and co-workers (7) to measure gross motor function in children with cerebral palsy (CP) or brain damage, enables changes in performance status to be evaluated after therapy or when monitored over time. We analysed its inter-rater and intra-rater reliability on the three most difficult dimensions. A video-recording of three children with CP performing test tasks was assessed on two occasions at an interval of six months by each of the 15 physiotherapists using the GMFM manual but without previous experience or training in the use of the instrument. Mean percentage scores were similar at the first and second assessments. Both inter- and intra-rater reliabilities were good, inter-rater reliability being 0.77 and 0.88 at the first and second assessments, respectively, and intra-rater reliability 0.68 at the second assessment. The findings suggest the GMFM to be a useful and reliable instrument for assessing motor function and treatment outcome in CP.

Cerebral Palsy↗

Reliability, validity, and sensitivity of a Swedish version of the revised and expanded Arthritis Impact Measurement Scales (AIMS2).

OBJECTIVE: To evaluate the reliability, validity, and sensitivity of a Swedish version of the Arthritis Impact Measurement Scales (AIMS2) in patients with rheumatoid arthritis (RA). METHODS: Reliability was assessed by a test-retest procedure with a 3-week interval and Cronbach's coefficient of internal consistency. Convergent validity was evaluated by correlation coefficients with the Health Assessment Questionnaire (HAQ), Mood Adjective Checklist (MACL), and disease activity variables. The sensitivity to change was assessed in 40 patients treated with disease modifying antirheumatic drugs, 24 mo after the start of the study. RESULTS: Significant differences were found for the Walking and Bending, Arthritis Pain, and Work scales (p < 0.05) in the test-retest reliability study. Internal consistency coefficients for the AIMS2 scales were 0.68-0.91. Convergent validity was established by significant correlations between the different scales in AIMS2 and the HAQ, MACL, and disease activity variables. Sensitivity to change was observed in 5 of 11 scales in AIMS2. CONCLUSION: After modifying some reliability problems due to a more differentiated scale and cultural differences between the Swedish and American patients with RA, the reliability, validity, and sensitivity to change of AIMS2 were found satisfactory.

Activities of Daily Living↗

Reliability and validity of stationary cystometry, stationary cysto-urethrometry and ambulatory cysto-urethro-vaginometry.

OBJECTIVE: The purpose of this article is to evaluate the reliability and validity of stationary cystometry, stationary cysto-urethrometry and ambulatory cysto-urethro-vaginometry for the diagnosis of urinary incontinence. MATERIAL AND METHODS: Literature search on reproducibility of cystometry is based on articles written in the English language found in Medline from 1987 to 1993. Five articles were found. These articles had references for another 6 older articles about the same topic. One recently published article is evaluated as well. Data about reproducibility of cystometry performed in these studies are given. Knowledge of reproducibility of cystometry combined with urethrometry, urethro-vagino or rectometry with leak detection during stationary or ambulatory recording is given. An evaluation of validity of the different methods is presented. RESULTS AND CONCLUSION: Conflicting data has been found for sensory parameters of stationary cystometry such as volume or pressure at first desire to void, strong desire to void and maximum cystometric capacity. No data were found concerning the reliability of cystometry for the recording of urinary incontinence. The main problem of urge-incontinent patients is leakage with a strong feeling of urgency. The recording of sensory parameters by cystometry is not valid when the main problem is urge incontinence. Cystometry will be valid for the diagnosis of urge incontinence when a leak detector is used or a mark is made at the tracings when leakage is observed. The reliability of cystometry concerning urge incontinence should be documented. The reliability of a urethral pressure recording when the urethral catheter is extracorporally fixed is low. The reliability and validity of ambulatory cysto-urethro-vaginometry in diagnosing urge incontinence, unstable urethra and genuine stress incontinence seems to be good, but needs better documentation.

Female↗

[Reliability and factorial structure of a rating scale for persistent vegetative state].

We developed a new rating scale, Kohnan Vegetative Score, to measure severity and small clinical changes in vegetative state patients. It has 7 items corresponding to the conditions of vegetative state by Japanese Society of Neurosurgery: motor function, food ingestion, urination and defecation, eye movement, vocalization, communication, and facial expression. Each item is rated in 5 ordinal categories: slight (score = 1), mild(2), moderate(4), and extreme(5). The sum of the scores is used as the summary score, which ranges from 7 to 35, and high score means 'severe'. We examined the reliability and the factorial structure of the Kohnan Vegetative Score. The subjects were 10 patients who met the conditions of vegetative state. Four neurosurgeons rated the subjects, and then 2 of them repeated the rating after one week interval. As a measure of reliability, the (weighted) Kappa coefficient proposed Cohen (1960, 1968) was calculated for each item, and the intraclass correlation coefficient (ICC) was calculated for the summary score. To analyze the factorial structure, the factor analysis was carried out. The minimum and the maximum weighted Kappa values were 0.44 and 0.64 for intra-rater reliability, and 0.37 and 0.69 for inter-rater reliability, respectively. Concerning the factorial structure, the contribution of the first factor was 91.5% which indicated the unidimensionality of the scale. The ICC's estimate for the summary score were 0.90 (95% C.I.: 0.766-0.970). On the basis of these results, the Kohnan Vegetative Score has unidimensionality and high reliability enough for a practical use.

Aged↗

An evaluation of the specificity, validity and reliability of jumping tests.

BACKGROUND: There were three objectives of this study: 1. To describe the influence of using a single and double leg take-off as a function of run-up length in jumping for height. 2. To determine if various types of jumps are specific in nature. 3. To evaluate two methods of assessing jumping height (a modified Vertec or Yardstick and a Board) for validity and inter-day reliability. METHODS: Seventeen male subjects were tested on jumps for height from a standing position and using a 1, 3, 5 and 7 stride run-up. These jumps were performed using a single and double leg take-off measured by the Yardstick. Selected jumps were also tested using a Board method and repeated for assessment of reliability. RESULTS: The single leg take-off produced significantly higher jumps when the run-up was three or more strides. The inter-relationships among jump conditions were generally high, however jump types could be considered as specific when the run-up length and number of legs used in the take-off were different. The Yardstick produced significantly greater jump heights than the Board method, which questions the validity of using a board for assessment of maximum jump performance. The reliability of both methods was generally high however the jumps performed from a run-up produced less reliable results than the standing jumps for the Yardstick. CONCLUSIONS: It was suggested that the design of tests to assess jumping ability should consider the specific jump type used in the sport of interest and that the Yardstick is the preferred mode of testing, provided that attempts are made to maximise reliability.

Adolescent↗

Reliability of closed double helix electrode for functional electrical stimulation.

The reliability of a closed double helix electrode in the lower limbs was studied. This electrode is an implanted intramuscular electrode and is used for a totally implantable functional electrical stimulation system. Eighty electrodes were evaluated retrospectively with a mean period of 15 months. The total implant time was 1222 electrode months. The cumulative proportion surviving was 0.934 at 6 months, 0.855 at 1 year, 0.765 at 2 years, and 0.730 after 30 months. Fifteen of 80 electrodes failed, seven showed increasing electrode impedance, and eight had undesirable changes in recruitment. Of the failed electrodes, 2/3 failed during the first 10 months. The reliability was 0.91 at 6 months and 0.80 at 1 year after implantation in all muscle groups. The closed double helix electrode displayed an increased reliability when compared with the open double helix electrode at 6 months, and an equivalent reliability as compared with the electrodes developed by Handa and colleagues at 6 months and 1 year, using the chi squared test for independence. This study suggests that the closed double helix electrode has an acceptable reliability and can be used as a part of a totally implantable functional electrical stimulation system.

Electric Stimulation Therapy↗

Reliability of manual skinfold tests in a healthy male population.

OBJECTIVE: To assess the intra- and interexaminer agreement of a manual skinfold thickness test and a manual skinfold compliance test. The relation between the weekly routine of the examiners and the intraexaminer reliability was also assessed for both tests. DESIGN: This is a reliability study of a common palpatory procedure to assess skinfold thickness and skinfold compliance. Twelve healthy subjects were palpated twice in two sessions by 12 examiners. SETTING: The study was conducted at the Polytechnic of Utrecht (the Netherlands), Faculty of Health Care, Department of Physiotherapy. SUBJECTS: Healthy male subjects recruited from students of the Polytechnic of Utrecht (the Netherlands), Department of Physiotherapy. RESULTS: The intraexaminer agreement Intraclass Correlation Coefficient [ICC(3.1)] was .25 for skinfold thickness and .28 for skinfold compliance. The interexaminer agreement [ICC(2,1)] ranged from .01 to .24. The Pearson correlation coefficient between the examiners age and routine vs. intraexaminer agreement ranged from -.41 to .23 (nonsignificant). CONCLUSIONS: The intra- and interexaminer agreement of the manual skinfold test produced poor-to-fair reliability. The correlation between the examiners' weekly routine and the intraexaminer reliability ranged from low negative to little (if any). This study shows a lack of reliability of palpatory tests for skinfold thickness and skinfold compliance. This outcome agrees with results of previous studies found in the literature.

Adult↗

[Reliability and validity studies with the triflexometer, a new method for assessing form and flexibility of the spine].

Developed by Orthotronic Medizintechnik GmbH, the so-called Triflexometer, version 3.22 was tested for its reliability and validity in measuring spinal posture and mobility. Reliability studies on 20 healthy subjects have shown this measurement method to be reliable, yet intra- and inter-rater reliability analyses also revealed that even for this healthy population discrepancies in the various measures may occur, both due to differences in compliance as well as fatigue and learning effects, and due to difficulties in stabilization of the normal posture, to a lesser extent due to certain specifics of the measurement technique (placing the markers, guiding the sensor). In total spinal immobility (ankylosing spondylitis), practically identical measurements are found, as is the case in dummy studies. The validity study on 20 healthy subjects found good correlations between the measurements obtaining using the triflexometer and those for double inclinometer, respectively, and that only minor mean value differences occur for the two methods. Also, triflexometer measurements for total anteflexion were found to correlate with those determined with the fingertip-to-floor method, no correlation was present however between the Triflexometer values and the Schober test. Triflexometer measurements performed on 114 healthy subjects of various ages served to prove that the range of spinal movement in the directions measured (sagittal and frontal) will reduce with age. To a lesser extent, this also applies to hip movement. Overall, our findings prove the triflexometer an easy-to-handle system which possesses high reliability and is suitable for valid and objective noninvasive assessment of global and segmental spinal mobility. Triflexometer examinations are highly uncomplicated to implement, and print-outs of the results obtained permit lasting documentation of the present status.

Adult↗

Temporomandibular disorders in children and adolescents: reliability of a questionnaire, clinical examination, and diagnosis.

Recently developed Research Diagnostic Criteria for Temporomandibular Disorders (RDC/TMD) have been shown to be reliable for diagnosing and assessing TMD in U.S. and Swedish adult populations; however, few studies have focused on clinical examination methods and diagnostic criteria for use with children and adolescents. The present study used a sample of 50 Swedish children and adolescents, aged 12 to 18 years, to evaluate usefulness and reliability of existing and specially developed measures and methods for assessing and diagnosing TMD in youth. Subjects underwent repeated clinical exams by two calibrated examiners to assess signs and symptoms per the RDC/TMD, and they responded to a specially developed self-administered questionnaire that addressed location and frequency of TMD-related pain and symptoms, jaw function, effect of pain on daily activities, and use of pain medications. Interexaminer and intraexaminer reliability was assessed for clinical examination, questionnaire items, and diagnosis. Reliability values ranged from acceptable to excellent for the RDC/TMD clinical exam and questionnaire, and from good to excellent reliability for measuring virtually all modified clinical parameters of TMD assessed in these young patients.

Adolescent↗

The Copenhagen Neck Functional Disability Scale: a study of reliability and validity.

OBJECTIVE: To determine whether a newly developed disability scale for patients with neck pain demonstrated acceptable reliability and validity. METHODS: Testing was conducted using three different samples of patients with neck pain (n = 162). Test-retest reliability of the scale was carried out on the same day with one sample (n = 39), and between-day reliability was carried out with another (n = 21). Differential item functioning with regard to the influence of gender and age was carried out with these two patient groups, as was construct validity. Responsiveness was measured using patients participating in a clinical trial involving patients with chronic neck pain (n = 102). Additionally, scale scores were compared with a wide range of physical measurements using the patients in the clinical trial. RESULTS: Short-term, between-day and postal questionnaire reliability coefficients were all extremely high. The Cronbach's alpha coefficient for internal consistency was 0.9 for the entire scale, and the coefficients for individual items were all greater than 0.88. Disability scale scores correlated strongly to pain scores as well as to doctor and patient global assessments, indicating good construct validity. Relative changes in disability scores demonstrated a moderately strong correlation to changes in pain scores after treatment. Scale scores correlated weakly to all physical measurements. CONCLUSIONS: The disability scale demonstrated excellent practicality and reliability. The scale accurately reflects patient perceptions regarding functional status and pain as well as doctor's global assessment and is responsive to change over long periods of time. We feel that this scale can be a valuable tool for the assessment of patients in future clinical trials and quality of care studies.

Adult↗

Reliability of scores on the Stroke Rehabilitation Assessment of Movement (STREAM) measure.

BACKGROUND AND PURPOSE: The Stroke Rehabilitation Assessment of Movement (STREAM) is a new clinical measurement tool for evaluating the recovery of voluntary movement and basic mobility following stroke. This article presents the results of 3 substudies examining the reliability (interrater and intrarater) and internal consistency of STREAM scores. SUBJECTS AND METHODS: A "direct-observation reliability study" was conducted on 20 patients who had strokes and were in a rehabilitation setting. Pairs of raters from a group of 6 participating therapists provided data to judge interrater agreement. A "videotaped assessments reliability study" was done to assess intrarater and interrater agreement on the scoring of videotaped performances using the STREAM measure and involved 4 videotaped assessments that were viewed and rated on 2 occasions by 20 physical therapists. The internal consistency of the STREAM scores was evaluated for 26 patients who had strokes and who demonstrated the full range of motor ability. RESULTS: The reliability of the STREAM scores was demonstrated by generalizability correlation coefficients of .99 for total scores and of .96 to .99 for subscale scores. The internal consistency of the STREAM scores was demonstrated by Cronbach alphas of greater than .98 on the subscales and overall. CONCLUSION AND DISCUSSION: These high levels of reliability support the use of the STREAM instrument for the measurement of motor recovery following stroke. Further work on the validity and responsiveness of the STREAM measure is in progress.

Aged↗

Validity and reliability in reporting sexual partners and condom use in a Swiss population survey.

OBJECTIVES: To examine the validity and reliability of indicators of sexual behaviour and condom use in annual telephone surveys (n=2800) of the general population aged 17 to 45 for the evaluation of AIDS prevention in Switzerland. METHODS: A test-retest study with additional focused interviews was conducted on a subsample (n=138) of the respondents aged 17 to 22 years. RESULTS: The subsample included more French speaking respondents (OR: 1.7, CI: 1.1-2.5) and more people in a stable relationship (OR: 2.2, CI: 1.5-3-3) than the initial sample but did not differ in any other way, although no data is available on their attitudes towards sex. The reliability of the indicators considered was high: number of lifetime, casual sex partners in the last 6 months and condom use with them, acquisition of a new steady partner during the year and condom use with this partner, condom use at last intercourse. However, the focused interviews raised questions about the validity of some of these indicators, presumably due to imprecise wording of the questionnaire items. Among sexually active respondents, 12.5% (95% CI: 4.7-25.5) of the men included non-penetrative sex in the definition of 'sexual intercourse', but only 1.9% (95% CI: 0.1-10.3) of the women. The propensity for men of counting acts or partners with whom no penetration had taken place in the total reported sex acts or partners was not significantly associated with any socio-demographic variables. In addition, among the 15 respondents who had reported consistent condom use with casual sex partners at interview, 40% (95% CI: 16.3-67.7) admitted at reinterview that sometimes they also had unprotected sex. CONCLUSIONS: The reliability of reports on sexual behaviour and condom use in this Swiss evaluation survey is good. The indicators derived from the annual surveys are robust measures and the monitoring of trends seems to be based on reliable measurement. However, more research is required on the validity of the data.

Acquired Immunodeficiency Syndrome↗

[The reliability of birth expectations in the Netherlands].

This study is concerned with the reliability of birth expectations in the Netherlands. The data are from four nationwide surveys: the National Survey on Fertility of 1969, the Netherlands Survey on Fertility and Parenthood Motivation of 1975, and the Netherlands Fertility Surveys of 1977 and 1982. These data permit the reliability of birth expectations data to be evaluated at the aggregate level only. Variations in the reliability of the data among the four surveys, and in the reliability of short- and long-term expectations, are analyzed. The results suggest that "asking for birth expectations in surveys in the Netherlands makes sense in so far that birth expectations can be used in...hypotheses for population forecasts." The precise ways of doing this are currently being studied at the Central Bureau of Statistics. (summary in ENG)

Attitude↗

The reliability of reporting of contraceptive behavior in DHS calendar data: evidence from Morocco.

This report addresses the consistency of reporting in the contraceptive calendar in the 1992 and 1995 Morocco Demographic and Health Surveys. Because a panel design was used in these surveys, the same women were interviewed in both years, providing a unique opportunity to examine the reliability of responses. Measures of reliability for various aspects of contraceptive-use dynamics are computed, and the impact of reporting errors on contraceptive failure, discontinuation, and switching rates is estimated. Results suggest that reporting of contraceptive behavior in Moroccan DHS calendar data appears to be relatively reliable at the aggregate level. Individual respondents, particularly those whose contraceptive patterns have been complex, have a lower level of reliability. The observed inconsistencies do not appear to affect aggregate-level estimates of contraceptive prevalence; however, measures of contraceptive-use dynamics are less stable.

Adult↗

Reliability and validity of NINCDS-ADRDA criteria for Alzheimer's disease. The National Institute of Mental Health Genetics Initiative.

OBJECTIVE: To assess interrater reliability and validity of NINCDS-ADRDA (National Institute of Neurological and Communicative Diseases and Stroke/Alzheimer's Disease and Related Disorders Association) criteria for Alzheimer's disease (AD). DESIGN: A multisite reliability and validity study in which clinicians from each site diagnosed 60 case summaries yielding a preconsensus estimate of reliability and validity. A consensus conference was conducted for each disagreement, leading to a postconsensus estimate of validity. The criterion standard was a diagnosis of AD by autopsy. SETTING: Three academic medical centers. SUBJECTS: A convenience sample of 60 detailed case summaries, 40 with AD and 20 with other dementing disorders. MAIN OUTCOME MEASURES: The kappa coefficient, sensitivity, and specificity. RESULTS: The kappa coefficient for preconsensus agreement on a diagnosis of probable or possible AD vs non-AD was 0.51; the sensitivity of a diagnosis of probable or possible AD for a pathological diagnosis of AD was 0.81, and the specificity was 0.73. The postconsensus sensitivity was 0.83, and the specificity was 0.84. CONCLUSIONS: The results support the reliability and validity of NINCDS-ADRDA criteria and show that the consensus process may improve diagnostic accuracy. The cases are reviewed with a focus on the sources of diagnostic disagreements and errors and possible changes that might improve the accuracy of the criteria.

Aged↗