Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

How young can children reliably and validly self-report their health-related quality of life?: an analysis of 8,591 children across age subgroups with the PedsQL 4.0 Generic Core Scales.

BACKGROUND: The last decade has evidenced a dramatic increase in the development and utilization of pediatric health-related quality of life (HRQOL) measures in an effort to improve pediatric patient health and well-being and determine the value of healthcare services. The emerging paradigm shift toward patient-reported outcomes (PROs) in clinical trials has provided the opportunity to further emphasize the value and essential need for pediatric patient self-reported outcomes measurement. Data from the PedsQL DatabaseSM were utilized to test the hypothesis that children as young as 5 years of age can reliably and validly report their HRQOL. METHODS: The sample analyzed represented child self-report age data on 8,591 children ages 5 to 16 years from the PedsQL 4.0 Generic Core Scales DatabaseSM. Participants were recruited from general pediatric clinics, subspecialty clinics, and hospitals in which children were being seen for well-child checks, mild acute illness, or chronic illness care (n = 2,603, 30.3%), and from a State Children's Health Insurance Program (SCHIP) in California (n = 5,988, 69.7%). RESULTS: Items on the PedsQL 4.0 Generic Core Scales had minimal missing responses for children as young as 5 years old, supporting feasibility. The majority of the child self-report scales across the age subgroups, including for children as young as 5 years, exceeded the minimum internal consistency reliability standard of 0.70 required for group comparisons, while the Total Scale Scores across the age subgroups approached or exceeded the reliability criterion of 0.90 recommended for analyzing individual patient scale scores. Construct validity was demonstrated utilizing the known groups approach. For each PedsQL scale and summary score, across age subgroups, including children as young as 5 years, healthy children demonstrated a statistically significant difference in HRQOL (better HRQOL) than children with a known chronic health condition, with most effect sizes in the medium to large effect size range. CONCLUSION: The results demonstrate that children as young as the 5 year old age subgroup can reliably and validly self-report their HRQOL when given the opportunity to do so with an age-appropriate instrument. These analyses are consistent with recent FDA guidelines which require instrument development and validation testing for children and adolescents within fairly narrow age groupings and which determine the lower age limit at which children can provide reliable and valid responses across age categories.

Adolescent↗

Reliability of follicle-stimulating hormone measurements in serum.

BACKGROUND: Follicle-stimulating hormone (FSH), a member of gonadotropin family, is critical for follicular maturation and ovarian steroidogenesis. Serum FSH levels are known to fluctuate during different phases of menstrual cycle in premenopausal women, and increase considerably after the menopause as a result of ovarian function cessation. There is little existing evidence to guide researchers in estimating the reliability of serum FSH measurements. The objective of this study was to assess the reliability of FSH measurement using stored sera from an ongoing prospective cohort--the NYU Women's Health Study. METHODS: Sixty healthy women (16 premenopausal, 44 postmenopausal), who donated at least two blood samples at approximately 1-year intervals were studied. An immunoradiometric assay using a sandwich monoclonal antibodies technique was used to measure FSH levels in serum. RESULTS: The reliability of a single log-transformed FSH measurement, as determined by the intraclass correlation coefficient, was 0.70 for postmenopausal women (95% confidence interval (CI), 0.55-0.82) and 0.09 for premenopausal women (95% CI, 0-0.54). CONCLUSIONS: These results suggest that a single measurement is sufficient to characterize the serum FSH level in postmenopausal women and could be a useful tool in epidemiological research. For premenopausal women, however, the reliability coefficient was low, suggesting that a single determination is insufficient to reliably estimate a woman's true average serum FSH level and repeated measurements are desirable.

Adult↗

Development and reliability of a self-report questionnaire to examine children's perceptions of the physical activity environment at home and in the neighbourhood.

BACKGROUND: Environmental factors are increasingly being implicated as key influences on children's physical activity. Few studies have comprehensively examined children's perceptions of their environment, and there is a paucity of literature on acceptable and reliable scales for measuring these. This study aimed to develop and test the acceptability and reliability of a scale which examined a broad range of environmental perceptions among children. METHODS: Based on constructs from ecological models, a survey incorporating items on children's perceptions of the physical and social environment at home and in the neighbourhood was developed. This was administered on two occasions, nine days apart, to a sample of 39 children aged 11 years (54% boys), attending a metropolitan Australian elementary school. The acceptability of the survey was determined by the proportion of missing responses to each item. The test-retest reliability of individual items, scores and scales were determined using Kappa statistics and percent agreement for categorical variables, and intraclass correlation coefficients (ICC) for continuous variables. RESULTS: There were few missing responses to each question, with only 4% of all responses missing. Although some Kappa values were low, all categorical variables showed acceptable reliability when examined for percent agreement between test and retest (range 68%-100% agreement). Continuous variables all showed moderate to good ICC values (range 0.72-0.92). CONCLUSION: Findings suggest this questionnaire is reliable and acceptable to children for assessing environmental perceptions relevant to physical activity among 11-year-old children.

Journal Article↗

Validation and test-retest reliability of the Royal Free Interview for Spiritual and Religious Beliefs when adapted to a Greek population.

BACKGROUND: The self-report version of the Royal Free Interview for Religious and Spiritual Beliefs has been confirmed as a valid and reliable scale, assessing the manner and nature in which spiritual beliefs are expressed. The aim of the present study was to evaluate the test-retest reliability and psychometric properties of the Greek version of the Royal Free Interview for Religious and Spiritual Beliefs. METHODS: A total of 209 persons (77 men and 132 women) with a mean age of 28.33 +/- 9.44 years participated in the study (test group). We subsequently approached 139 participants of the test group with a mean age of 28.93 +/- 9.60 years, who were asked to complete the Royal Free Questionnaire a second time two weeks later (retest group). RESULTS: The vast majority of participants (58.9%) reported both a religious and a spiritual belief, compared to 52 (25.1%) who told of a religious belief only. The internal consistency of the spiritual scale for the test group proved to be good, as standardized inter-item reliability / Cronbach's alpha was 0.83. Item-total correlations ranged from 0.51 to 0.73. They indicated very good levels of differentiation, thus showing that the questions were appropriate. Internal consistency of the spiritual scale for the retest group proved as good as for the test group. Standardized inter-item reliability / Cronbach's alpha was 0.84. Item-total correlations ranged from 0.52 to 0.75. The Pearson correlation coefficient for the total test-retest score of the spiritual scale was 0.754 (p < 0.001). CONCLUSION: The Greek version of the Royal Free Interview for Religious and Spiritual Beliefs is reliable and thus suitable for use in Greece.

Journal Article↗

The Footwear Assessment Form: a reliable clinical tool to assess footwear characteristics of relevance to postural stability in older adults.

OBJECTIVE: Falls in older adults are common and may result in serious injury. Inappropriate footwear has been suggested to be a contributing factor to many falls. However no studies have been undertaken to determine whether clinicians can reliably assess footwear variables thought to influence postural stability in older adults. The aim of this study was therefore to develop a simple clinical footwear assessment form and assess its reliability, both between examiners and with repeated assessments over time. DESIGN: Two examiners assessed seven footwear variables (shoe type, heel height, heel counter stiffness, longitudinal sole rigidity, sole flexion point, tread pattern and sole hardness) in 12 different shoes, and repeated the measurements three weeks later. The examiners were blinded to each other's and their own previous results. RESULTS: Analysis using the kappa (kappa) and percentage agreement statistics revealed the examiners' footwear assessments to be generally highly reliable (kappa = 0.47-1.00 for inter-tester comparisons, kappa = 0.40-1.00 for intra-tester comparisons), with the exception of inter-tester assessment of sole hardness (kappa = 0.03-0.48). CONCLUSION: The Footwear Assessment Form is a reliable clinical tool for the assessment of shoe type, heel height, heel counter stiffness, longitudinal sole rigidity and tread pattern; however, a more objective protocol may be required to improve the reliability of sole hardness evaluation. The Footwear Assessment Form can now be used with confidence in the clinical setting and in future investigations to determine the contribution of footwear characteristics to instability and falls in older adults.

Accidental Falls↗

A study to compare the reliability of composite finger flexion with goniometry for measurement of range of motion in the hand.

OBJECTIVE: To establish the intra- and inter-rater reliability of composite finger flexion (CFF), and to compare this with goniometry. DESIGN: Fifty-one physiotherapists and occupational therapists took part in the study. The hand of a normal subject was splinted in three different positions. Using a goniometer and a ruler alternately, each therapist measured both the proximal interphalangeal joint and CFF of three digits, following a standardized protocol. This process was repeated three times. SETTING: Eighteen NHS hospital sites in the UK. RESULTS: The two measurement methods produced different ranges and standard deviations for each digit. The repeatability coefficient shows that repeated intra-rater goniometric measures fall within 4-5 degrees of each other 95% of the time. Inter-rater goniometric measures fall within 7-9 degrees. Repeated intra-rater CFF measures fall within 5-6 mm of each other, whereas inter-rater fall within 7-9 mm. The influence of occupation, experience in hand therapy, years of practice and routine use were found to have no effect on reliability. Scaling of the two methods of measurement allowed comparison between them to be made. CFF and goniometry are equally reliable when comparing inter-rater reliability, but goniometry displays less variability than composite finger flexion for intra-rater measurements. CONCLUSION: In this study involving a subject with normal joints, goniometry is more reliable than CFF when only one measurer is involved. However, CFF may be a useful alternative where multiple joint measures are required, or when goniometry is impracticable.

Analysis of Variance↗

Good inter-rater reliability of the Frenchay Activities Index in stroke patients.

OBJECTIVE: To assess inter-rater reliability of the Frenchay Activities Index (FAI) when used by occupational therapists in stroke patients. DESIGN: Independent administration of the FAI by occupational therapists in 45 stroke outpatients. SETTING: Outpatient departments of two Dutch rehabilitation centres. STATISTICS: Agreement between pairs of raters was assessed at item level using weighted Kappa and the percentage absolute agreement and for the total score using intra-class correlations (ICC), Bland-Altman plots and computation of the smallest real difference (SMR). RESULTS: ICC of the total score was good (0.90; 95% confidence interval (95% CI): 0.82-0.94). The Bland-Altman plot did not reveal relationships between agreement and height of scores. Reliability at item level was good (kappa >0.60) in 11 out of 15 items. The item 'gainful work' showed high agreement but a low kappa due to an extremely skewed score distribution. Reliability was insufficient in three items: 'local shopping', 'social occasions' and 'actively pursuing hobby'. Clarification of scoring instructions of these items might further improve the inter-rater reliability of the FAI. CONCLUSION: The FAI is a reliable instrument to measure outcomes of outpatient rehabilitation in patients with stroke.

Activities of Daily Living↗

The intra-rater reliability of the balance performance monitor when measuring sitting symmetry and weight-shift activity after stroke in a community setting.

OBJECTIVE: To examine the intra-rater reliability of sitting symmetry and weight-shift activity measurements in poststroke adults. DESIGN: An intra-rater reliability study. SETTING: A community setting. SUBJECTS: Adult stroke survivors attending stroke support groups within the community of Nottingham (U.K.). MAIN MEASURES: The Balance Performance Monitor used to measure sitting symmetry and weight-shift activity. Intraclass correlation coefficients (ICCs) and their 95% confidence intervals (95% CI) were calculated. The Bland Altman method for assessing agreement is also presented. RESULTS: We tested 49 participants (median age 73 years; interquartile range 68-81 years). Between-test reliability for sitting symmetry was high: ICC (1,1) = 0.93 (95% CI 0.87 < or = ICC < or = 0.96). The mean difference between the measures (d) was -0.08 (95% CI -0.48 < or = d < or = 0.31); the standard deviation of the differences (SDdiff) was 1.383. The coefficient of repeatability was 2.76; the 95% limits of agreement were -2.850 and 2.682. Between-test reliability for weight-shift activity was also high: ICC (1,1) = 0.86 (95% CI 0.77 < or = ICC < or = 0.92). Bland-Altman d = -0.08 (95% CI -0.19 < or = d < or = 0.35), SDdiff = 0.936. The coefficient of repeatability was 1.87; the 95% limits of agreement were -1.792 and 1.952. CONCLUSIONS: The 95% CI for d for both parameters crossed zero, indicating that between-test bias is unlikely. Sitting symmetry and weight-shift activity measures demonstrated acceptable levels of reliability.

Aged↗

Electrical stimulation of human tibialis anterior: (A) contractile properties are stable over a range of submaximal voltages; (B) high- and low-frequency fatigue are inducible and reliably assessable at submaximal voltages.

OBJECTIVES: To investigate the validity and reliability of submaximal voltage stimulation for assessing the 'fresh' contractile properties of human tibialis anterior muscle (TA) and the efficacy of such stimulation in inducing and assessing high- and low-frequency fatigue. INTERVENTIONS: (A) Contractile properties of fresh TA were assessed in six normal volunteers using multifrequency stimulation trains (comprising 2 seconds at each of 10, 20 and 50 Hz, arranged contiguously) over a range of submaximal voltages. (B) On three separate occasions, fatigue was induced in the TA of 10 normal volunteers by means of a 3-minute unbroken sequence of the described multifrequency stimulation trains, delivered at a 'standardized' submaximal voltage. This fatiguing protocol was preceded by discrete multifrequency stimulation trains, at the same standardized voltage, but followed by discrete multifrequency trains delivered over a range of submaximal voltages (which included the standardized voltage). OUTCOME MEASURES: In experiment A the 10:50 Hz and 20:50 Hz force ratios were analysed for between-voltages variability using coefficients of variation (CVs), and for trends using Friedman tests and post-hoc Wilcoxon tests. In experiment B low-frequency fatigue was detected using 10:50 Hz and 20:50 Hz force ratios derived from the discrete multifrequency trains. High-frequency fatigue was calculated from the decline in high-frequency force which occurred during the fatiguing protocol itself. Each parameter was assessed for between-days repeatability using CVs. RESULTS: In experiment A the 'fresh' 10:50 Hz force ratio was clearly unreliable at voltages which generated <10% of maximal voluntary contractile force (MVC) (CV< or =29.7%), but was reasonably reliable at voltages which generated 20-30% of MVC (CV < or = 11.5%; p = 0.847). The 'fresh' 20:50 Hz force ratio was,in contrast, extremely reliable throughout the tested voltage range (CV< or =5.8%; p = 0.636) in fresh muscle. In experiment B paired t-tests indicated that the fatiguing protocol induced significant high-frequency fatigue (p <0.0037) and low-frequency fatigue (p <0.0008 for 'fresh' versus 'fatigued' 10:50 Hz force ratio; p <0.0001 for 'fresh' versus 'fatigued' 20:50 Hz force ratio). In muscle thus fatigued, the 20:50 Hz force ratio was extremely reliable in the 20-33% of MVC range (CV < or =7.3%; p = 0.847). Between-days repeatability was poor for the 10:50 Hz force ratio in both fresh and fatigued muscle (CV < or =23.8 and 44.4% respectively), but was highly acceptable for both voluntary and stimulated fatigue indices and for the 20:50 Hz force ratio, the latter in both fresh and fatigued muscle. CONCLUSIONS: These results confirm the validity and reliability of submaximal voltages in assessing contractile properties (including low-frequency fatiguability) and inducing fatigue of human TA.

Adult↗

Validity and reliability of the MSQLI in cognitively impaired patients with multiple sclerosis.

Multiple sclerosis (MS) has important effects on quality of life but it is unknown how cognitive impairment affects the ability to assess or report this. Our objective was to determine whether cognitive impairment negatively affects the construct validity and the reliability of the Multiple Sclerosis Quality of Life Inventory (MSQLI). A neuropsychological test battery and the Multiple Sclerosis Functional Composite (MSFC) were administered to a sample of 136 patients referred for cognitive testing by their neurologists. Age, sex, education and ethnicity-adjusted T scores were calculated for each cognitive variable. Cognitive impairment was defined as any T score less than the fifth percentile. The MSQLI was administered prior to neuropsychological testing and readministered one to four weeks later: Correlations between the MSFC and the SF-36 were determined and compared between the cognitively impaired and unimpaired groups as the main test of construct validity. Test-retest and internal consistency reliability of each of the scales were compared for the impaired and unimpaired groups. Seventy-six (56%) patients were cognitively impaired. Construct validity and internal consistency reliability did not differ between the cognitively impaired and unimpaired groups. Test retest reliability was lower for the bladder and vision scales in the impaired group, but remained acceptable for the bladder scale (r > 0.7). Cognitive impairment, a common MS manifestation, does not appear to reduce the reliability or validity of the MSQLI as a patient self-report measure of health status and quality of life.

Adult↗

Reliability and validity of pinch and thumb strength measurements in de Quervain's disease.

The purpose of this study is to evaluate the test-retest reliability and construct validity of pinch and thumb strength measurements in subjects with de Quervain's disease. Maximal palmar pinch and thumb strength (adduction, extension, abduction, and flexion) were measured using, respectively, a pinch gauge and a biaxial dynamometer. The reliability was estimated using the generalizability theory. The validity hypotheses were as follows: 1) the pinch and thumb strength of the symptomatic side would be significantly lower than that of the asymptomatic side, and 2) the strength loss would be greater for thumb extension and abduction. The reliability was high for all strength measurements, pinch strength being the more reliable one. The pinch and thumb strength in all directions evaluated was significantly decreased on the symptomatic side (p<0.003); no direction showed a greater decrease than the others. The results suggest that pinch and thumb strength measurements are reliable and able to show a decreased strength on the symptomatic side in this population.

Adult↗

A method for automatic identification of reliable heart rates calculated from ECG and PPG waveforms.

OBJECTIVE: The development and application of data-driven decision-support systems for medical triage, diagnostics, and prognostics pose special requirements on physiologic data. In particular, that data are reliable in order to produce meaningful results. The authors describe a method that automatically estimates the reliability of reference heart rates (HRr) derived from electrocardiogram (ECG) waveforms and photoplethysmogram (PPG) waveforms recorded by vital-signs monitors. The reliability is quantitatively expressed through a quality index (QI) for each HRr. DESIGN: The proposed method estimates the reliability of heart rates from vital-signs monitors by (1) assessing the quality of the ECG and PPG waveforms, (2) separately computing heart rates from these waveforms, and (3) concisely combining this information into a QI that considers the physical redundancy of the signal sources and independence of heart rate calculations. The assessment of the waveforms is performed by a Support Vector Machine classifier and the independent computation of heart rate from the waveforms is performed by an adaptive peak identification technique, termed ADAPIT, which is designed to filter out motion-induced noise. RESULTS: The authors evaluated the method against 158 randomly selected data samples of trauma patients collected during helicopter transport, each sample consisting of 7-second ECG and PPG waveform segments and their associated HRr. They compared the results of the algorithm against manual analysis performed by human experts and found that in 92% of the cases, the algorithm either matches or is more conservative than the human's QI qualification. In the remaining 8% of the cases, the algorithm infers a less conservative QI, though in most cases this was because of algorithm/human disagreement over ambiguous waveform quality. If these ambiguous waveforms were relabeled, the misclassification rate would drop from 8% to 3%. CONCLUSION: This method provides a robust approach for automatically assessing the reliability of large quantities of heart rate data and the waveforms from which they are derived.

Algorithms↗

Reliability and validity of functional neuroimaging techniques for identifying language-critical areas in children and adults.

Advances in neuroimaging technologies over the last 15 years have prompted their relatively widespread use in the study of brain mechanisms supporting language function in children and adults. We reviewed reliability and external validity studies of 3 of the most common functional imaging methods, functional magnetic resonance imaging (fMRI), magnetoencephalography (MEG), and positron emission tomography (PET). Although reliability and validity reports for fMRI are generally quite favorable, significant variability was found across studies with respect to methodology, preventing in some cases either the assessment of the reliability of individual datasets, or cross-study comparisons. Reliability and validity reports of MEG are strong, yet methodological questions regarding optimal modeling techniques remain. PET investigators report good concordance of language maps with data from more invasive brain mapping techniques, but its use of radioactive tracers and poorer spatial and temporal resolution make it the least optimal of the 3 methods for language mapping. Investigations of the cortical networks supporting language function during development and into adulthood should be viewed in the context of the validity and reliability of the methods used, with careful attention to details regarding the methodologies employed in the acquisition and analysis of statistical maps.

Aging↗

Reliability and validity of dyspnea measures in patients with obstructive lung disease.

Dyspnea, the clinical term for shortness of breath, is the primary symptom and an important outcome measure in evaluations of patients with lung disease. It is a subjective symptom that has proved difficult to quantify. Many dyspnea measures are available, yet it is difficult, based on the existing literature, to determine the most reliable and valid. In this study, we evaluated 6 measures of dyspnea for reliability and validity: (a) Baseline Dyspnea Index (BDI) and Transition Dyspnea Index, (b) UCSD Shortness of Breath Questionnaire (SOBQ),(c) American Thoracic Society Dyspnea Scale, (d) Oxygen Cost Diagram, (e) Visual Analog Scale, and (f) Borg Scale. Subjects were 143 patients (74 women) and 69 men) with obstructive lung disease, ages 40 to 86, FEV(1.0) 0.36 to 3.53 L, FVC 1.07 to 5.74 L. Dyspnea measures were assessed for test-retest reliability internal consistency, interrater reliability, and construct validity (i.e., correlations among dyspnea measures and correlations of dyspnea measures with exercise tolerance, health-related quality of life, lung function, anxiety, and depression). Results suggest that the SOBQ and BDI demonstrated the highest levels of reliability and validity among the dyspnea measures examined.

Journal Article↗

Further studies of a measure of adience-abience: reliability.

Previous studies have suggested adequate reliability for a measure of perceptual adience-abience based on the HABGT in view of significant evidence concerning its construct and predictive validity. The present study explored the test-retest reliability of this scale with 40 process schizophrenics over a two-week interval. Reliability was found to be adequate for both males and females (rho equals .84) for the total scale. The four components of the scale were similarly found to be reliable. No subject changed on retest with respect to adient or abient orientation. Interjudge reliability was very high (rho = .912).

Adult↗

Psychopathology scale of the Hutt adaptation of the Bender-Gestalt Test: reliability.

The test-retest reliability of the Hutt Adaptation of the Bender-Gestalt test was explored with a population of 40 process schizophrenics over a two-week interval. The total Psychopathology Scale Score was found to have high retest reliability for both male and female patients (rho = .87 for males and .83 for females). Moreover the three major components for the Scale were found to have high reliability, and fairly high reliabilities were obtained for patients scoring high as well as low on the Scale. Interjudge reliability was also found to be very high (rho = .895), confirming previous studies in this respect. On these grounds, the Scale offers promise both for clinical and research purposes.

Adult↗

Assessing the person reliability of an individual MMPI protocol.

This study investigated various measures commonly employed to assess the person reliability of an individual Minnesota Multiphasic Personality Inventory (MMPB protocol. Specifically, relationships among indices of person reliability and the standard MMPI validity scales were examined using the responses of 82 subjects who completed the MMPI on two occasions separated by 1 week. Person reliability indices were based on within-occasion responses to identical and to psychologically similar items, and on three across-occasion response consistency measures. The validity scales, namely, the L, F, K, and Cannot Say scales, showed higher test-retest stability than the within-occasion person reliability indices. Further, the validity scales and person reliability indices appeared to reflect multiple facets of dependable responding. Interestingly, an individual's tendency to change responses to MMPI items from the test to the retest was significantly predictable. Clinical implications of these findings were derived.

Journal Article↗

Reliability of the MMPI-2 Subtle and Obvious scales with psychiatric inpatients.

Internal consistency reliability estimates were determined for the Wiener and Harmon (1948) Subtle and Obvious scales of the Minnesota Multiphasic Personality Inventory-2 (MMPI-2) with a group of psychiatric inpatients. Except for Subtle Scale 3 (Hy), Subtle scale reliabilities were unacceptably low, ranging from .32 to .59. In each case, inclusion of the Subtle items attenuated the reliability of the corresponding Full clinical scale. Removing the Subtle items resulted in comparable reliabilities across all Minnesota Multiphasic Personality Inventory-2 (MMPI-2) scales, ranging from .75 to .92. It was hypothesized that the poor reliability of the Subtle scales may attenuate the validity of the Full clinical scales.

Adjustment Disorders↗