Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Effort-limited treadmill walk test: reliability and validity in subjects with postpolio syndrome.

OBJECTIVE: To determine the reliability and construct validity of an effort-limited treadmill walk test to measure functional ability in subjects with postpolio syndrome in an outpatient postpolio clinic. DESIGN: Functioning and distance walked on a treadmill to a Borg "hard" effort level were measured three times, a week apart, by two blinded raters in 15 subjects with postpolio syndrome, aged 37-67 yrs, with new weakness, fatigue, and pain but with no other cause of symptomatology or condition-limiting walking. One rater tested them twice. Fatigue activity level, mobility, and health-related quality of life (Medical Outcome Study Short Form Health Survey [SF-36]) defined functioning. Generalizability correlation coefficients determined intrarater, test-retest and interrater reliability. The correlations relating the distance walked and functioning determined construct validity. RESULTS: Reliability for generalizability correlation coefficients were: intrarater, 0.91; test-retest, 0.85; and interrater, 0.58. Interrater reliability improved to 0.91 with adherence to a standardized protocol. Validity was established with correlations between the distance walked and SF-36 physical component score (0.66), physical role (0.60), bodily pain (0.60), and vitality (0.55). CONCLUSIONS: The treadmill walk test provides a reproducible and valid measure of ability in persons with postpolio syndrome with a single rater, but a standardized protocol is essential for reliability.

Adult↗

Reliability and validity of the GMFM-66 in 0- to 3-year-old children with cerebral palsy.

OBJECTIVE: To examine the reliability and validity of a 66-item version of the Gross Motor Function Measure (GMFM-66) to assess the gross motor functions of children <3 yrs old with cerebral palsy. DESIGN: 298 valid samples were obtained from 171 children with cerebral palsy (male 126, female 45, mean age 19 mos, age range 3-36 mos) measured with GMFM-88. Then a 73-item version of GMFM (GMFM-73) special for these children was obtained by Rasch analysis. GMFM-66 score and GMFM-73 scores of each sample were obtained. The reliability and validity of GMFM-66 were evaluated by analyzing the correlation between the scores and between the change scores of these two GMFM versions. The relative precision of GMFM-73 vs. GMFM-66 was also analyzed. RESULTS: The test-retest reliability and interscorer reliability of GMFM-66 were both high. The ICC scores were 0.9666 and 0.9782, respectively. Significant correlations were found between the scores (r = 0.9848, P < 0.001) and between the change scores (r = 0.8700, P < 0.001) of these two versions of GMFM. A 14% less gain in relative precision was achieved when using GMFM-73 vs. GMFM-66. CONCLUSION: Results indicated that the GMFM-66 had good reliability and validity in assessing the gross motor functions of children <3 yrs old with cerebral palsy. The GMFM-73 derived in the present study did not function significantly better for young children than GMFM-66.

Cerebral Palsy↗

Three types of skin-surface thermometers: a comparison of reliability, validity, and responsiveness.

OBJECTIVE: To compare the reliability, validity, and responsiveness of a thermistor thermometer (thermistor) and two different infrared thermometers (one designed to measure tympanic temperature and one for skin temperature). DESIGN: Reliability and validity were evaluated by making two separate measurements from the skin at identical spots of each hand, forearm, shoulder, thigh, shin, and foot in 17 healthy subjects. Intramuscular temperature was recorded at the hand and shin sites. Test-retest reliability was calculated using intraclass correlation for each instrument. Pearson correlation assessed the relationship between the skin and intramuscular temperatures at the hand and shin sites (validity). Each instrument's ability to measure temperature change (responsiveness) was assessed by measuring skin temperatures serially from 17 limbs of ten patients with complex regional pain syndrome undergoing intravenous regional sympathetic blockade. Responsiveness index values were calculated. RESULTS: Reliability was strong and similar for each device (intraclass correlation: thermistor = 0.96, tympanic = 0.96, skin = 0.97), as was validity (r: thermistor = 0.90, tympanic = 0.92, skin = 0.92). Responsiveness was marginally better for the infrared skin device (responsiveness index: skin = 4.2, tympanic = 3.6, thermistor = 3.6). CONCLUSIONS: For the purposes of clinical electrodiagnostic laboratory and other physiatry applications, the performance of the infrared thermometers is equal to or superior to that of the traditionally used thermistor. All three devices are highly reliable and valid, whereas the infrared skin device is slightly more responsive. Infrared thermometers have the advantage of being quicker to operate and more portable.

Adult↗

The patient and observer scar assessment scale: a reliable and feasible tool for scar evaluation.

At present, various scar assessment scales are available, but not one has been shown to be reliable, consistent, feasible, and valid at the same time. Furthermore, the existing scar assessment scales appear to attach little weight to the opinion of the patient. The newly developed Patient and Observer Scar Assessment Scale consists of two numeric scales: the Patient Scar Assessment Scale (patient scale) and the Observer Scar Assessment Scale (observer scale). The patient and observer scales have to be completed by the patient and the observer, respectively. The patient scale's consistency and the observer scale's consistency, reliability, and feasibility were tested. For the Vancouver Scar Scale, which is the most frequently used scar assessment scale at present, the same statistical measurements were examined and the results of the observer scale and the Vancouver scale were compared. The concurrent validity of the observer scale was tested with a correlation to the Vancouver scale. Furthermore, the authors examined which specific characteristics significantly influence the general opinion of the patient and the observers on the scar areas. Four independent observers have each used the observer scale and the Vancouver scale to assess 49 burn scar areas of 3 x 3 cm belonging to 20 different patients. Subsequently, the patients completed the patient scale for their scar areas. The (internal) consistency of both the patient and the observer scales was acceptable (Cronbach's alpha, 0.76 and 0.69, respectively), whereas the consistency of the Vancouver scale appeared not to be acceptable (alpha, 0.49). The reliability of the observer scale completed by a single observer was acceptable (r = 0.73). The reliability of the Vancouver scale completed by a single observer was lower (r = 0.69). The observer scale showed better agreement than the Vancouver scale because the coefficient of variation was lower (18 percent and 22 percent, respectively). The concurrent validity of the observer scale in relation to the Vancouver scale is high (r = 0.89, p < 0.001). Linear regression of the general opinions on scars of the observer and the patient showed that the observer's opinion is influenced by vascularization, thickness, pigmentation, and relief, whereas the patient's opinion is mainly influenced by itching and the thickness of the scar. Such an impact of itching and thickness of the scar on the patient's opinion is an important and novel finding. The Patient and Observer Scar Assessment Scale offers a suitable, reliable, and complete scar evaluation tool.

Adolescent↗

The validity and reliability of chinese frontal assessment battery in evaluating executive dysfunction among Chinese patients with small subcortical infarct.

OBJECTIVES: Frontal Assessment Battery (FAB) is a valid and reliable screening test for evaluating executive dysfunction among whites with frontal and subcortical degenerative lesions. We studied the properties of a Chinese version of FAB (CFAB) in evaluating executive dysfunction among Chinese stroke patients with small subcortical infarct. METHODS: Concurrent validity was evaluated using Wisconsin Card Sorting Tst (WCST) and Mattis Dementia Rating Scale-Initiation/Perseveration Subset (MDRS I/P) among 41 controls and 30 stroke patients with small subcortical infarct. Discriminant validities of CFAB and its subitems were compared with those of Mini-Mental State Examination (MMSE). Internal consistency, test-retest, and interrater reliability of CFAB were evaluated. RESULTS: The CFAB had low to good correlation with various executive measures: MDRS I/P (r = 0.63, p < 0.001), number of category completed (r = 0.45, p < 0.001), and number of perseverative errors (r = -0.37, p < 0.01) of WCST. Among the executive measures, only number of category completed had significant but small contribution (6.5%, p = 0.001) to the variance of CFAB. A short version of CFAB using three items yielded higher overall classification accuracy (86.6%) than that of CFAB full version (80.6%) and MMSE (77.6%). Internal consistency (alpha = 0.77), test-retest reliability (rho = 0.89, p < 0.001), and interrater reliability (rho = 0.85, p < 0.001) of CFAB were good. CONCLUSION: Although CFAB is reliable, it is only moderately valid in evaluating executive dysfunction among Chinese stroke patients with small subcortical infarct. The clinical use of CFAB in the evaluation of executive dysfunction among this group of patients cannot be recommended at this stage.

Aged↗

The reliability and concurrent validity of the figure-of-eight method of measuring hand edema in patients with burns.

OBJECTIVE: Water volumetry is considered the "gold standard" for hand edema assessment. This technique requires considerable time, staff, and specialized equipment. The figure-of-eight method for hand edema assessment has been tested only in the orthopedic population. The objective of this study was to test the reliability and concurrent validity of the figure-of-eight method of measuring hand edema in the burn patient population. METHODS: We conducted a prospective blinded study with 20 burned patients (33 edematous hands) admitted from February to May 2005. Two testers performed three separate blinded measurements on each edematous hand, using the figure-of-eight technique. A third tester performed two measurements, using water volumetry. An independent investigator recorded all measurements. Intratester and intertester reliability were analyzed. Concurrent validity was examined and compared with water volumetry measurements. RESULTS: Intraclass correlation coefficients (ICC) for the intratester reliability of the figure-of-eight method were 0.96 for tester 1 and 0.97 for tester 2. The ICC for intertester reliability of the figure-of-eight measurements was 0.94. The intratester ICC for volumetric measurements was 0.99. Correlation coefficient (Pearson's) for tester 1 was 0.83 (P < .01), and for tester 2, 0.89 (P < .01). CONCLUSION: The figure-of-eight technique is a reliable and valid measurement tool for measuring hand edema. This technique is a more clinically feasible tool than water volumetry in the burn patient population.

Adult↗

Chaos-induced modulation of reliability boosts output firing rate in downstream cortical areas.

The reproducibility of neural spike train responses to an identical stimulus across different presentations (trials) has been studied extensively. Reliability, the degree of reproducibility of spike trains, was found to depend in part on the amplitude and frequency content of the stimulus [J. Hunter and J. Milton, J. Neurophysiol. 90, 387 (2003)]. The responses across different trials can sometimes be interpreted as the response of an ensemble of similar neurons to a single stimulus presentation. How does the reliability of the activity of neural ensembles affect information transmission between different cortical areas? We studied a model neural system consisting of two ensembles of neurons with Hodgkin-Huxley-type channels. The first ensemble was driven by an injected sinusoidal current that oscillated in the gamma-frequency range (40 Hz) and its output spike trains in turn drove the second ensemble by fast excitatory synaptic potentials with short term depression. We determined the relationship between the reliability of the first ensemble and the response of the second ensemble. In our paradigm the neurons in the first ensemble were initially in a chaotic state with unreliable and imprecise spike trains. The neurons became entrained to the oscillation and responded reliably when the stimulus power was increased by less than 10%. The firing rate of the first ensemble increased by 30%, whereas that of the second ensemble could increase by an order of magnitude. We also determined the response of the second ensemble when its input spike trains, which had non-Poisson statistics, were replaced by an equivalent ensemble of Poisson spike trains. The resulting output spike trains were significantly different from the original response, as assessed by the metric introduced by Victor and Purpura [J. Neurophysiol. 76, 1310 (1996)]. These results are a proof of principle that weak temporal modulations in the power of gamma-frequency oscillations in a given cortical area can strongly affect firing rate responses downstream by way of reliability in spite of rather modest changes in firing rate in the originating area.

Action Potentials↗

Fast unambiguous stereo matching using reliability-based dynamic programming.

An efficient unambiguous stereo matching technique is presented in this paper. Our main contribution is to introduce a new reliability measure to dynamic programming approaches in general. For stereo vision application, the reliability of a proposed match on a scanline is defined as the cost difference between the globally best disparity assignment that includes the match and the globally best assignment that does not include the match. A reliability-based dynamic programming algorithm is derived accordingly, which can selectively assign disparities to pixels when the corresponding reliabilities exceed a given threshold. The experimental results show that the new approach can produce dense (> 70 percent of the unoccluded pixels) and reliable (error rate < 0.5 percent) matches efficiently (< 0.2 sec on a 2GHz P4) for the four Middlebury stereo data sets.

Algorithms↗

Reliability of the next of kins' estimates of critically ill patients' quality of life.

The aim of this study was to determine the reliability and validity of relatives' assessment of patients' quality of life and to measure the agreement between patients' and relatives' responses to the Short Form 36 quality of life questionnaire, at discharge from and 6 months following intensive care treatment. Ninety-nine patient-relative pairs were studied. Reliability was quantified by using measures of internal consistency (Cronbach's alpha and correlation coefficients) and reliability coefficients. Relatives' responses met the required standards of reliability and validity, but reliability was consistently weaker in the mental health dimension. Relatives' and patients' scores differed significantly in six dimensions at discharge; however, agreement between patients' and relatives' responses, as measured by the Kappa statistic, was fair, improved over 6 months, and was greatest in aspects concerning physical health. We conclude that relatives are able to give a good proxy assessment of functional aspects of quality of life.

Adolescent↗

Interobserver reliability between a nurse and anaesthetist of tests used for predicting difficult tracheal intubation.

We examined the interobserver reliability, between a nurse and anaesthetist, of five tests used to predict difficult tracheal intubation: mouth opening; thyromental distance; head and neck movement; mandibular luxation; and assessment of oropharyngeal view. For each test, an anaesthetic nurse and a specialist registrar anaesthetist were trained to use a standard method of examination. Most of the tests had either good or very good reliability. Assessment of mouth opening demonstrated only moderate reliability and assessment of oropharyngeal view demonstrated poor reliability. The interobserver reliability estimates between a nurse and an anaesthetist are similar to those previously demonstrated between two anaesthetists.

Head Movements↗

Reliability and validity of the Greek version of Kogan's Old People Scale.

AIMS AND OBJECTIVES: The aim of this study was to test the psychometric properties--validity and reliability--of the Greek version of Kogan's Old People scale. BACKGROUND: The ageing of the population in most of the developed world and in Greece, challenge-nursing care, therefore, nursing education needs to be updated accordingly. Until today there have been no studies in Greece in relation to student nurses' attitudes towards older people. To have a reliable questionnaire for measuring a Greek population's attitudes towards older people the Kogan's Old People Scale was translated into Greek. DESIGN: The study was designed as a cross-sectional survey. The main reason for choosing a cross-sectional survey was the time limits for the study. A sample of 390 nursing students in Athens participated in the study. METHODS: A questionnaire was given to the students, which included the Kogan's Old People Scale. RESULTS: Results showed Cronbach's alpha coefficient 0.73 for the OP- scale and 0.65 for the OP+ scale, which are comparable to published studies until today. The six-factor solution explains the 41.5% of the variance in the sample. The scale was also found to differentiate between first and final year students on how their education in nursing is influencing their attitudes towards the older people. CONCLUSIONS: Reliability and validity supported the Greek version of the Kogan's Old People Scale as a reliable instrument. Its use in evaluating Greek nursing education programmes could help in preparing nurses capable of meeting the needs of older people. RELEVANCE TO CLINICAL PRACTICE: Nursing education--basic and lifelong--needs to be updated in order to respond to the needs of older people and a reliable instrument can help to evaluate it.

Aged↗

Reliability and validity of a maternal food frequency questionnaire designed to estimate consumption of common food allergens.

BACKGROUND: Maternal food intake during pregnancy may influence the development of food hypersensitivity (FHS) in the child. A food frequency questionnaire estimating the frequency with which some of the mains food allergens are consumed was designed and validated. MATERIALS AND METHODS: Pregnant women were recruited at the ante-natal clinic of St. Mary's Hospital, Isle of Wight, UK. A food frequency questionnaire was developed and validated by comparing responses to information recorded in 7 days food diaries. The reliability of the food frequency questionnaire was evaluated by asking women to complete the questionnaire on two separate occasions at 30 and 36 weeks gestation. RESULTS: Fifty-seven women completed the validity study and 91 women completed the reliability study. For both validity and reliability, questions with dichotomous response categories showed the highest level of agreement. Frequency of intake of foods commonly "hidden" in foods produced the lowest validity and reliability scores. In the validity study responses to the food frequency questionnaire identically matched information recorded in the food diaries 80% of the time, on average. In the reliability study, responses were identical on both questionnaires 85% of the time on average. CONCLUSION: In this study a food frequency questionnaire estimating the frequency with which some of the main food allergens are consumed during pregnancy was designed and validated. This food frequency questionnaire could be used in future studies to assess the role of maternal food intake in the development of FHS in the infant.

Adolescent↗

The psychometric properties of the Social Training and Achievement Record: II. Reliabilities and concurrent validities.

Three studies are reported which investigated the test-retest and inter-rater reliabilities of the Social Training and Achievement Record as well as concurrent validities with intelligence. Reliabilities of total and sub-scale scores as well as individual items are reported. Test-retest reliabilities over the period of one month were good for total scores, subscale scores and individual items. Inter-rater reliabilities were more modest. At times, inter-rater reliabilities of a small proportion of items were only at chance levels. Concurrent validities with Raven's Matrices and the Weschler Adult Intelligence Scale were comparable with other measures of adaptive behaviour. The role of direct observation in the assessment of adaptive behaviour, in particular with respect to individual skills being selected for teaching, is discussed.

Activities of Daily Living↗

Are caregivers' reports of motivation valid? Reliability and validity of the Reiss Profile MR/DD.

BACKGROUND: Sensitivity theory proposes that there are wide individual differences in what motivates people with intellectual disability. The Reiss Profile MR/DD is a rating scale that measures 15 fundamental motives. This study examined the internal consistency and interrater reliability of the 15 subscales as well as the validity of motivational profiles. METHOD: The study consisted of two distinct but related steps. First, the interrater reliability of the rating scale was established by having pairs of raters evaluate 48 individuals. Second, raters were presented with three different motivational profiles and asked to identify which one corresponded to the individual they had rated 4 weeks earlier. RESULTS: Results indicated good internal consistency (average alpha=0.84), significant variability in the interrater reliability (average intraclass correlation coefficient=0.52), and excellent validity (95% of the correct profiles were chosen). Average discrepancies between pairs of raters are presented. CONCLUSIONS: Interrater reliability is an important topic for professionals working in the field of intellectual disability and results are discussed in terms of the factors that affect it. This is the first published study to report on the interrater reliability of the Reiss Profile MR/DD.

Adaptation, Psychological↗

A new approach to define defect extensions of endodontically treated teeth: inter- and intra-examiner reliability.

For endodontically treated teeth, there are no standardized measures available to define the extent of loss in tooth substance prior to final restoration. In this study, defect size was classified and the applicability of the classification was tested related to the inter- and intra-examiner reliability. For classification, three parameters were investigated: (i) remaining tooth substance in the vertical dimension (level A-D, aspect I), (ii) remaining tooth substance as regarded horizontally (mm; bucco-lingual and mesio-distal, aspect II), and (iii) size of the orifice (mm; aspect III). Four non-calibrated or (pre-trained) examiners were asked to gauge and classify 20 casts of clinically broken down teeth. The measurements were repeated twice every alternative week giving three separate readings. Inter-examiner reliability was determined at weeks 1, 3 and 5. The intra-examiner reliability was compared between readings 1 and 2, 1 and 3, and 2 and 3. As statistical tests, intra-class correlation (ICC) and Cohen's kappa (weighted) were used at a significance level of P < 0.05. Inter- and intra-examiner reliability for ordinal data (aspect I) revealed, with one exception, 'moderate' to 'very good' evaluations. Inter- and intra-examiner reliability (ICC) of metric data of aspect II and III was primarily 'excellent'. It may be concluded that the newly developed classification could be applied as an appropriate and reproducible method to define defect extension in endodontically treated teeth.

Humans↗

Reliability, validity and efficiency of multiple choice question and patient management problem item formats in assessment of clinical competence.

Despite a lack of face validity, there continues to be heavy reliance on objective paper-and-pencil measures of clinical competence. Among these measures, the most common item formats are patient management problems (PMPs) and three types of multiple choice questions (MCQs): one-best-answer (A-types); matching questions (M-types); and multiple true/false questions (X-types). The purpose of this study is to compare the reliability, validity and efficiency of these item formats with particular focus on whether MCQs and PMPs measure different aspects of clinical competence. Analyses revealed reliabilities of 0.72 or better for all item formats; the MCQ formats were most reliable. Similarly, efficiency analyses (reliability per unit of testing time) demonstrated the superiority of MCQs. Evidence for validity obtained through correlations of both programme directors' ratings and criterion group membership with item format scores also favoured MCQs. More important, however, is whether MCQs and PMPs measure the same or different aspects of clinical competence. Regression analyses of the scores on the validity measures (programme directors' ratings and criterion group membership) indicated that MCQs and PMPs seem to be measuring predominantly the same thing. MCQs contribute a small unique variance component over and above PMPs, while PMPs make the smallest unique contribution. As a whole, these results indicate that MCQs are more efficient, reliable and valid than PMPs.

Clinical Competence↗

Reliability and learning from the objective structured clinical examination.

The difficulties in measurement of the clinical performance of students in the health professions are well known by educators. One innovative measure incorporated in several of the educational programmes, including the BSc in Nursing programme, in the Faculty of Health Sciences, at McMaster University, Hamilton, Ontario, Canada is the objective structured clinical examination (OSCE). The purpose of this study was to determine the reliability of this evaluation method, both within and between stations. One problem that has been noted by users of the OSCE method is that performance on individual OSCE stations is poorly correlated across stations, apparently regardless of the particular content of the station. A number of hypotheses have been advanced to attempt to explain this phenomenon: performance of any skill is sufficiently variable that the correlation is poor; different skills have little common basis, so that there is no generalizability from one to another, or reliability of assessment in any one station is low. To test these hypotheses, a study was designed for test-retest and interrater reliability. Students undergoing a 10-station OSCE also repeated their starting OSCE station at the end of the examination circuit. In addition, several stations were rated by more than one observer (interrater). This study of 71 first-year BScN students showed that the interrater reliability was high (ICC = 0.80 to 0.99), and test-retest reliability on the same station was good (ICC = 0.66 to 0.86); however, correlation across stations was low (alpha = 0.198). Thus it is apparent that there is high consistency of repeated performance of a skill but little consistency of performance on different skills.

Clinical Competence↗

Reliability and efficiency of components of clinical competence assessed with five performance-based examinations using standardized patients.

The present study was conducted to provide in-depth information on the reliabilities of measures of the separate components of clinical competence (e.g. data collection, test interpretation, diagnosis, etc.) assessed by a performance-based examination consisting of standardized-patient cases administered to five classes of senior medical students at Southern Illinois University School of Medicine. In general, the reliabilities of the competencies as they were actually measured on the examination (using the number of cases on which each competency was actually measured) were small, with 54% less than 0.30 and 75% less than 0.40. For generalizability coefficients pooled across the five classes and projected to a common number of k = 10 cases, two of the nine competencies had reliabilities in the 40s, a third was close to 0.40, and the remaining six were in the low 20s. The number of cases needed for the competencies to achieve the recommended reliability of 0.80 ranged from 45 to 170 cases, with six of the nine competencies requiring over 100 cases to reach the 0.80 level. The low reliabilities of these measures of the components of clinical competence raise serious questions about using the scores as indicators of student performance.

Clinical Clerkship↗