Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Evaluation of the reliability of reflective marker placements.

This study determined the reliability of recording knee angle coordinates in the sagittal plane during erect stance by use of skin reflective markers. To this end, 32 healthy male and female subjects participated in three standardised test sessions of 1 hour duration spaced at 1- and 4-week intervals. At each session, markers were placed on three bony landmarks of the leg and their coordinates were photographed with the leg in extension and semi-flexion. The resultant knee angles, calculated from lines drawn on the photographs marking the coordinate sites, were subjected to separate repeated measures analysis of variance statistical procedures. The results showed that both sets of measurements could be carried out reliably when repeated within a single test session (R = 0.87). However, when repeated over time, neither measurement reached the criterion for acceptable reliability (set at R > or = 0.80). The magnitude of the method error was also higher between test sessions for both angles than it was for tests repeated within the sessions. These data suggest that whereas knee coordinates recorded on a single day can be reproduced reliably, they may be less useful for evaluating knee angle changes over time.

Adult↗

Measurements of scapular position and rotation: a reliability study.

Smooth motion of the scapula and humerus with respect to the thorax is essential for shoulder function and abnormalities may indicate clinical entities. Recent studies have made an attempt to devise simple, practical means of quantifying scapular position. The aim of this study was to examine the intra-tester and inter-tester reliability of two methods and to determine if significant differences existed between the dominant versus non-dominant extremity. Seventeen healthy volunteers (4 M; 13 F) were examined by two testers. The tape measurements consisted of the classic methods of Kibler and DiVeta in three sitting postures, expanded by the measurement of the linear distance from the medial border to the thoracic mid-line, and the scapular size measure. The SAS software package was used for data analysis. The Intraclass Correlation Coefficient (ICC) intra-tester reliability ranged between 0.96-0.8 for both methods without significant differences, whereas the ICC for inter-tester reliability ranged between 0.42-0.9 with higher values (moderate and good) for the Kibler technique. In the additional tests high values were also obtained for ICC intra-tester, except for the measurements of the linear distance of the medial border of the scapula to the thoracic mid-line and the distance of the inferior process of the acromion to the third vertebra, both in 90 degrees abduction and internal rotation. The ICC for inter-tester was only acceptable for the DiVeta measurement on 45 degrees abduction. Significant differences were noted between both testers on the following measures: Kibler in 45 degrees abduction, DiVeta in 45 degrees abduction and 90 degrees abduction and the scapular size measure. The comparison of dominant versus non-dominant extremity revealed larger but not significantly different means for the dominant extremity in the classic methods. Significant differences occurred for Tester 1 in the measurement of the distance of the medial border to the thoracic mid-line and Tester 2 in DiVita in 45 degrees abduction. The SEM values never exceeded 1 cm. We believe that the Kibler technique holds promise for further studies, has the advantage of measuring in three positions and with some familiarisation can be reliable. Further research is necessary in patients with pathological conditions.

Adult↗

Assessing intrarater, interrater and test-retest reliability of continuous measurements.

In this paper we review the problem of defining and estimating intrarater, interrater and test-retest reliability of continuous measurements. We argue that the usual notion of product-moment correlation is well adapted in a test-retest situation, whereas the concept of intraclass correlation should be used for intrarater and interrater reliability. The key difference between these two approaches is the treatment of systematic error, which is often due to a learning effect for test-retest data. We also consider the reliability of a sum and a difference of variables and illustrate the effects on components. Further, we compare these approaches of reliability with the concept of limits of agreement proposed by Bland and Altman (for evaluating the agreement between two methods of clinical measurements) and show how product-moment correlation is related to it. We then propose new kinds of limits of agreement which are related to intraclass correlation. A test battery to study the development of neuro-motor functions in children and adolescents illustrates our purpose throughout the paper.

Child↗

Measuring interrater reliability among multiple raters: an example of methods for nominal data.

This paper reviews and critiques various approaches to the measurement of reliability among multiple raters in the case of nominal data. We consider measurement of the overall reliability of a group of raters (using kappa-like statistics) as well as the reliability of individual raters with respect to a group. We introduce modifications of previously published estimators appropriate for measurement of reliability in the case of stratified sampling frames and we interpret these measures in view of standard errors computed using the jackknife. Analyses of a set of 48 anaesthesia case histories in which 42 anaesthesiologists independently rated the appropriateness of care on a nominal scale serve as an example.

Reproducibility of Results↗

The interobserver reliability of off-line antral follicle counts made from stored three-dimensional ultrasound data: a comparative study of different measurement techniques.

OBJECTIVE: To assess the interobserver reliability of antral follicle counts (AFCs) made from stored three-dimensional (3D) ultrasound data using conventional two-dimensional (2D) images, 3D multiplanar view and 3D-rendered 'inversion mode'. METHODS: 3D transvaginal ultrasound was performed in the early follicular phase (days 2-5) of the menstrual cycle in 41 subjects aged < 40 years, undergoing investigation for subfertility. From the stored 3D ultrasound datasets, the number of antral follicles of 2-10 mm in diameter in each ovary was independently measured, using all three methods by three investigators, each with a different level of experience. The image quality of each dataset was subjectively categorized into one of three groups, based on the proportion of the ovarian contour that could be seen clearly. RESULTS: There was no significant difference in the mean AFC between the observers for any of the three different techniques. The intraclass correlation coefficient (ICC) for the 2D-equivalent mode, the 3D multiplanar mode, and the 3D-rendered inversion mode were indicative of good interobserver reliability for each method. The interobserver reliability for the 3D-rendered inversion mode was better with Grade 1 image quality than with Grade 3 image quality. There were no equivalent differences, however, between the three different grades of image quality with the 2D-equivalent and 3D multiplanar modes. The time taken for AFC measurement using 3D-rendered inversion mode was significantly longer than with the 2D equivalent and 3D multiplanar methods. CONCLUSIONS: 3D image displays and rendering techniques do not appear to offer any advantage over a conventional 2D display in terms of AFC measurement reliability. AFC measurement using the 3D-rendered inversion mode has an adequate interobserver reproducibility but is dependent on image quality.

Adolescent↗

Reliability and validity of tissue volume measurement by three-dimensional ultrasound: an experimental model.

OBJECTIVE: To determine the validity and the intra- and interobserver reliability of volume measurements of an endometrium-like model using a three-dimensional (3D) ultrasound rotational technique. METHODS: A 3D ultrasound dataset was obtained from a sample of bovine liver containing a portion of chicken chest muscle (CCM). The process was repeated seven times using pieces of CCM of different sizes, resulting in seven datasets. Each portion of CCM was then placed in a water-filled volume-scaled tube and the 'actual' volumes were calculated by water displacement. For each dataset, ten volumes were calculated by each of two observers using a (VOCAL) with a 15 degrees rotational step. Reliability was assessed by calculating intraclass correlation coefficients (ICC) and validity by examining the percentage difference from the actual volume using limits of agreement. RESULTS: The volume measurement of organic tissues using the 3D ultrasound rotational method was highly reliable (intraobserver ICC, 0.998 for Observer 1 and 0.997 for Observer 2; interobserver ICC, 0.997) and valid (the bias and 95% limits of agreement of the percentage difference from the actual volume was only 0.57 (-3.07 to 4.21) % for Observer 1 and - 0.17 (-4.34 to 4.0) % for Observer 2). CONCLUSIONS: The 3D sonographic measurement, using VOCAL with a 15 degrees rotational step, of small and irregular tissues is reliable and valid, suggesting that it is a useful technique for measurement of the endometrial volume and other volumes of similar size.

Animals↗

On the reliability of implicit and explicit memory measures.

Functional dissociations between implicit and explicit memory tests often take the form of large differences between groups or experimental conditions (e.g., amnesics and controls, elderly and younger persons, or persons learning with and without a distracting secondary task) when performance is assessed using explicit memory tests, whereas no difference is observed with implicit memory tests. We argue that the interpretation of such dissociations in terms of the memory processes or systems involved in performance is problematic because the same data pattern would emerge as a result of a mere methodological artifact, that is, the situation that implicit memory tests have low reliability whereas explicit memory tests are fairly reliable measurement instruments. We present reasons for such a reliability difference, and we demonstrate it empirically in Experiments 1a, 1b, and 2. However, our analysis also shows, and Experiment 3 confirms empirically, that implicit memory tests need not necessarily be less reliable measurement instruments than explicit memory tests.

Adolescent↗

Reliability concept as a trend in biophysics of aging.

The vitality of a living system is determined by the reliability characteristics of its functional elements at different organizational levels--from enzymes up to the organism as a whole. The new field of biophysics, in dealing with the problem of reliability, incorporates theoretical and experimental studies on quantitative characteristics and mechanisms of failure and renewal processes in biological systems. It also includes the elaboration of methods for testing the reliability and predicting the failures in biological systems. Apart from the formal fitting to the mortality data, the theory of reliability can serve as an investigative approach for searching for realistic mechanisms of aging.

Aging↗

Are sample sizes usually at least an order of magnitude too low for reliable estimates of leaf asymmetry?

Estimates of leaf size and asymmetry for individual trees are often obtained using sample sizes that are too small to take into account the possibility that size and asymmetry may be affected by the position of the leaf on the tree. This issue was addressed by exploring variation in leaf size and asymmetry within an individual of Alder (Alnus glutinosa). We found differences between branches for leaf size and for signed asymmetry but not for unsigned asymmetry. We also found that the size of a leaf was not correlated with its position on a branch and that the asymmetry of a leaf was not correlated with either its position on a branch or with the asymmetry of its neighbour. Repeated subsampling of a sample of 870 leaves showed that a subsample size approaching 500 leaves was required for consistently reliable estimates of the standard deviation of unsigned asymmetry. Smaller subsamples were required for consistently reliable estimates of mean unsigned asymmetry and of the mean and standard deviation of leaf size, but subsamples of less than 100 leaves provided consistently reliable estimates only of mean leaf size. For this species, reliable estimates of an individual's level of asymmetry are obtained only if several hundred leaves are sampled over several branches, but it is not necessary to sample the same sequence of leaves from each branch.

Analysis of Variance↗

A reliability test of standard-based quantitative PCR: exogenous vs endogenous standards.

The quantitative measurement of gene expression requires consistent and reliable standards. At least two categories of standards, endogenous and exogenous, are currently used for quantitative PCR. The reliability of these two methods, however, has not been carefully compared. We hypothesized that a reliable quantitative PCR assay would be able to detect known dilutions of a given single-stranded (ss-) cDNA. By measuring VEGF ss-cDNA copy numbers or signal ratios of GAPDH/VEGF in 10x and 100x diluted samples of two original ss-cDNA preparations, an exogenous recombinant DNA standard (a VEGF-mimic plasmid) and an endogenously expressed GAPDH standard were tested for their ability to detect dilution factors. Using the recombinant DNA standard, the dilution factor was detected as 10.3 and 135.0 in 10x and 100x diluted samples of the original CaSki cell ss-cDNA, respectively. The detected dilution factors were 12.3 and 226.2, respectively, in 10x and 100x diluted ss-cDNA from U-251 MG cells. On the other hand, with the endogenous GAPDH standard, the dilution factors were detected as 2.7 and 8.0 in the same 10x and 100x dilutions of the original U-251 MG cell ss-cDNA. Using the same endogenous GAPDH standard, the detected dilution factors were both 4.8 in 10x and 100x dilutions of the original CaSki cell ss-cDNA. It was also found that the number of endogenous copies of GAPDH mRNA was about 1000 times higher than VEGF. The high internal lockup ratio of GAPDH vs VEGF copy numbers and the requirement for additional primer pairs make the use of an abundant endogenous standard an unreliable choice in quantitative or semi-quantitative PCR. In contrast, exogenous standard-based quantitative PCR was shown to be an accurate and reliable method for the quantitation of gene expression.

Central Nervous System Neoplasms↗

Torque, total work, power, torque acceleration energy and acceleration time assessed on a dynamometer: reliability of knee and elbow extensor and flexor strength measurements.

Isometric torque and isokinetic peak torque, total work, power, torque acceleration energy and acceleration time at 30, 120 and 240 degrees.s-1 of the knee and elbow extensors and flexors were measured using an isokinetic dynamometer in 24 healthy women. Intrasession variation of the measurements was evaluated and the short-term and long-term reliability was assessed by repeating all procedures after averages of 2 and 30 days, respectively. The effect of learning on peak torque during a session was also evaluated. Moreover, the effect of general warming-up on knee extensor and flexor strength was examined on a separate day. Using correlations, numerous studies have indicated that muscle strength measurements are reliable. Correlations, however, are inappropriate and misleading in studies on reliability. In the present study reliability of each strength variable was expressed as the coefficient of variation (CV). With the protocol used, neither learning nor warming-up had any significant effect on strength. As expected, intra-session variation tended to be less than short-term and long-term inter-session variation. The CVs for strength variables measured 30 days apart exceeded 5% for all variables and rose to 107% for acceleration time. Substantial between-subject variation of individual CVs were found. The study demonstrated that muscle strength measurements may be highly unreliable in the individual subject.

Acceleration↗

The reliability of the hole-board apparatus.

Two aspects of the reliability of the hole-board apparatus were investigated-the similarity between scores of different samples of the same population on their first exposure to the apparatus, and the test-retest reliability. Rats and mice were given a 5-min exposure to the hole-board and then retested for 5 min after 1, 2 or 8 days. Male rats and mice showed good initial exposure reliability, whereas the female mouse groups differed significantly. All animals showed a positive test-retest correlation (range 0.31-0.78), but a homogeneous group (e.g. all animals habituating) produced higher correlations (range 0.60-0.99). Comparison of scores on the two 5-min exposures showed that not all groups showed significant habituation, but the animals exposed to the hole-board for two 10-min periods showed both significant habituation and test-retest reliability.

Animals↗

Reliability of the behavior problem checklist with institutionalized male delinquents.

Interrater and 2-week test-retest reliability coefficients were determined for subscales of the Behavior Problem Checklist on 50 males incarcerated in a state receiving facility for delinquent adolescents. Raters were 22 dormitory counselors, 2 of whom rated each child after 1 week and again after 3 weeks of observing the boys. Interrater reliability ranged from .06 to .68 on the various BPC subscales and was .50 overall, reflecting wide variation in the agreement of raters in rating boys on different dimensions. Test-retest reliability coefficients for the same rater at 2-week intervals were higher (.71 overall) and also varied among subscales. Raters were able to agree best on aggressive, acting-out behaviors. Other personality dimensions tapped by the BPC were rated with less reliability in this particular setting.

Adolescent↗

Reliability and validity studies of endoluminal ultrasonography for anorectal disorders.

PURPOSE: Endoluminal ultrasonography (ELUS) is accurate in the assessment of penetration through the rectal wall by carcinoma. Clinical studies were performed to determine the reliability and validity of ELUS. METHODS: The interobserver reliability among four observers with varying experience with ELUS was determined for staging the penetration of rectal cancer through the rectal wall. The ability of ELUS to change the clinical management of the referring clinician (comprehensiveness) was assessed on all referrals over a six-month period. RESULTS: The reliability of ELUS for staging rectal cancer demonstrated only fair to moderate correlation (weighted kappa range, 0.22-0.47). The accuracy of ELUS compared with surgical pathology demonstrated a learning curve proportional to the experience of the observer. In 45 percent of referrals, ELUS changed the clinical management of patients and in 76 percent of referrals the clinician's confidence in the diagnosis and management of patients was altered. ELUS was more likely to change the management of patients with pelvic pouch sepsis (70 percent) and early neoplastic lesions (57 percent) than in more advanced neoplastic lesions (40 percent), perianal Crohn's disease (40 percent), complex noninflammatory bowel disease sepsis (33 percent), and incontinence (31 percent). CONCLUSIONS: ELUS has the ability to change the clinical management of a variety of anorectal conditions. However, for neoplasia the interobserver reliability is only moderate and a learning curve exists.

Adenocarcinoma↗

The reliability of the SADS-LA in a family study setting.

The joint-rater and test-retest reliability study of two translated versions of the SADS-LA (Schedule for Affective Disorders and Schizophrenia--Lifetime version--modified for the study of anxiety disorders), one in French and the other in German, have been tested in family study settings, in a sample of patients and first-degree relatives. The test-retest reliability study demonstrated that identification of major affective disorders and schizophrenia was performed with sufficient reliability; however, diagnoses of subtypes of major disorders (e.g. bipolar II disorder) and identification of minor disorders was less reliable. The implications of these findings in phenotype identification during family studies in psychiatry are discussed.

Adult↗

Diet measurement in Vietnamese youth: concurrent reliability of a self-administered food frequency questionnaire.

Dietary patterns of Asian Americans change with increasing acculturation, leading to increased consumption of Western foods including those high in fat. Strategies to preserve the healthy aspects of traditional diets need to be developed and dietary assessment methods evaluated. Little is known about reliability of brief dietary measures in the general population or among minority youth. The concurrent reliability of a brief food frequency questionnaire (FFQ) was determined among Vietnamese youth using diet reports. Students in a bilingual high school program were given a FFQ. Students then completed daily diet reports one day each week over seven weeks. The data from the FFQ were compared to the daily food reports. The reliability of the FFQ was highest for frequently eaten food types like rice (r = 0.626, P < 0.01), fruit (r = 0.513, P < 0.01), meat (r = 0.525, P < 0.01) and vegetables (r = 0.474, P < 0.01) and was lower for less commonly eaten types including fish/shellfish (r = 0.227, P = 0.20) and fried foods (r = 0.310, P = 0.07). These results suggest that a few simple FFQ items, particularly for indicator foods such as rice, are reliable for dietary assessment in this population.

Adolescent↗

Validity and reliability of undergraduate performance assessments in an anesthesia simulator.

PURPOSE: To examine the validity and reliability of performance assessment of undergraduate students using the anesthesia simulator as an evaluation tool. METHODS: After ethics approval and informed consent, 135 final year medical students and 5 elective students participated in a videotaped simulator scenario with a Link-Med Patient Simulator (CAE-Link Corporation). Scenarios were based on published educational objectives of the undergraduate curriculum in anesthesia at the University of Toronto. During the simulator sessions, faculty followed a script guiding student interaction with the mannequin. Two faculty independently viewed and evaluated each videotaped performance with a 25-point criterion-based checklist. Means and standard deviations of simulator-based marks were determined and compared with clinical and written evaluations received during the rotation. Internal consistency of the evaluation protocol was determined using inter-item and item-total correlations and correlations of specific simulator items to existing methods of evaluation. RESULTS: Mean reliability estimates for single and average paired assessments were 0.77 and 0.86 respectively. Means of simulator scores were low and there was minimal correlation between the checklist and clinical marks (r = 0.13), checklist and written marks (r = 0.19) and clinical and written marks (r = 0.23). Inter-item and item-total correlations varied widely and correlation between simulator items and existing evaluation tools was low. CONCLUSIONS: Simulator checklist scoring demonstrated acceptable reliability. Low correlation between different methods of evaluation may reflect reliability problems with the written and clinical marks, or that different aspects are being tested. The performance assessment demonstrated low internal consistency and further work is required.

Anesthesiology↗

[Reliability of home monitoring with event-recording compared with polysomnography in infants].

AIM OF THE STUDY: Although reliable recognition of hypoxemia, apnea, bradycardia and tachycardia is absolutely necessary for home monitoring, many commercially available home monitors have not been sufficiently tested for their sensitivity. The purpose of this study was to determine the reliability of the home monitor VitaGuard 3000 by comparing it with the manual evaluation of full polysomnography (polysomnographic system Alice 3). PATIENT AND METHODS: 20 infants (11 males; 9 females) aged between 5 and 40 weeks (12.2 +/- 3.5 weeks, median 10.5) and with a gestational age between 29 and 41 weeks (37.7 +/- 6.1 weeks, median 39.0) were tested using both full polysomnography and, simultaneously, the home monitor VitaGuard 3000. The results were evaluated manually and compared. RESULTS: The monitor system detected 7/51 central apneas, 6/260 desaturations and 7/18 tachycardias. The sensitivity was 13.72% for apnea, 4.23% for desaturation and 38.80% for tachycardia. The reliability of the home monitor for detecting apnea, desaturation and tachycardia was therefore insufficient. CONCLUSION: A polysomnographic reliability test should be mandatory for all home monitoring systems prior to commercial introduction.

Age Factors↗