Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Lifetime DSM-IV diagnosis of alcohol, cannabis, cocaine and opiate dependence: six-month reliability in a multi-site clinical sample.

Psychiatric research increasingly emphasizes the diagnosis of symptoms and syndromes on a longitudinal basis. This study tests the reliability of lifetime DSM-IV diagnoses of alcohol, cannabis, cocaine and opiate dependence. The CIDI-SAM was administered at intervals not less than six months apart to a multi-site sample of 201 clinical respondents. The reliability of lifetime diagnosis of the syndromes, of the criteria which constitute the syndromes, and of the ages of onset reported for the criteria and for the dependence syndromes as a whole, were studied and the effects of patient characteristics suspected to degrade reliability were examined. There was generally good agreement, statistically, at both the syndrome and criterion level between the two interviews. Lifetime diagnoses for three of the drugs--alcohol, cannabis and opiates--were made at or near levels of agreement generally considered excellent under less strict testing conditions, and cocaine dependence was only marginally below this level. Most criteria showed good reliability and all delivered about equal results when averaged across the four substances, although a relationship between reliability and centrality of the symptom to the individual drug abuse pattern was found. Age of onset was almost uniformly highly reliable. Most patient characteristics bore no detectable relationship to reliability, although patients with multiple drug use patterns may warrant more careful probing by interviewers. Overall, these data indicate that lifetime symptoms and diagnoses can be queried reliably, although they must be reported with less confidence than current state diagnoses.

Age of Onset↗

Determinants of reliability in psychiatric surveys of children aged 6-12.

The reliability of young children's self reports of psychiatric information is a concern of epidemiologists and clinicians alike. This paper explores the determinants of test-retest reliability in a sample of children from the general population using reliability coefficients constructed from a kappa statistic. Age, cognitive ability, and gender are related to consistency of reports in a test-retest paradigm. Controlling for age, cognitive ability and gender, children report more reliably on observable behaviors, and less reliably on questions involving unspecified time, reflections of one's own thoughts, and comparison of themselves with others. The reliability of reports of emotions lies between these two extremes. Surprisingly, sentence length of up to 40 words and psychiatric impairment of the child as measured by the Child Global Assessment Scale did not influence reliability. As might be expected, parents' reports of their children are more reliable than their children's reports.

Affective Symptoms↗

Effects of cognitive impairment on the reliability of geriatric assessments in nursing homes.

OBJECTIVE: To explore the relationship between an elderly subject's cognitive status and the reliability of multidimensional assessment data. DESIGN: Survey, with cognitive status as the independent variable and interrater reliability as dependent variable. SETTING: Medicare/Medicaid-certified nursing homes. PARTICIPANTS: 147 residents age 65 or older. MEASUREMENTS: Dual assessments of elderly nursing home residents were performed by nurse assessors using the Health Care Financing Administration's new Minimum Data Set for Nursing Home Resident Assessment and Care Screening (MDS). Assessments were classified on the basis of residents' cognitive status, and levels of disagreement between assessors were analyzed. MAIN RESULTS: Overall assessment reliability, agreement concerning a resident's activities of daily living status, and the reliability of estimates of his or her communication skills and sensory abilities were significantly affected by a resident's cognitive status. The presence of cognitive impairment made these measurements less reliable--especially those related to communication skills, vision, and hearing. CONCLUSIONS: Assessments of residents suffering from cognitive impairment were significantly less reliable than assessments of cognitively intact residents. However, these differences in reliability were not uniform across all assessment domains. When treating the cognitively impaired elderly, clinicians must exercise caution in their reliance on standardized measurements that may be less reliable for this population.

Activities of Daily Living↗

Reliability of a standardized and expanded Brief Psychiatric Rating Scale: a replication study.

This study aimed to determine the replicability of the interrater reliability coefficients obtained with a standardized and expanded Brief Psychiatric Rating Scale (BPRS-E) in a 1991 psychometric evaluation. Furthermore, intrarater reliability was assessed. At item level, interrater concordance turned out to be satisfactory for most of the BPRS-E items. However, only a few of the items reached acceptable chance-corrected coefficients. In contrast to the previous study, the anxiety-depression subscale met the standard of acceptable interrater reliability in the present study. As in the 1991 study, the 10-item psychotic disintegration scale as well as BPRS-18 global scores met (or closely approximated) this standard. The 6 additional items of BPRS-E did not contribute to the scale's reliability. Joining the samples of the 1991 and replication studies (to cover the range of symptoms' severity and heterogeneity more fully) did not improve interrater reliability. Intrarater reliability coefficients were globally comparable to interrater reliability coefficients. In all, the results of this replication study suggest that only the anxiety-depression subscale, the 10-item psychotic disintegration scale and the BPRS-18 global scale can be used reliably in unselected groups of psychiatric inpatients in acute distress.

Adolescent↗

Factors affecting reliability coefficients of health attitude scales.

This study determined the minimum number of health attitude items and minimum sample size required to achieve maximum scale reliability coefficients, using different methods of estimating reliability. A 54-item alcohol attitude scale was administered to 700 participants. The scale produced .96 and .91 reliability coefficients, using the Cronbach Alpha (CA) and the Split-half (S-B) methods, respectively. A computer program randomly selected groups of participants and items from the pool of participants and items using different increments. A matrix of coefficients of reliability for both methods was calculated for different groups of items and sample size. To replicate the study, a 30-item cancer attitude scale was administered to more than 1,000 representative participants and produced reliability coefficients of .94 (using CA) and .82 (using S-B). The same computer and statistical procedures were repeated for the second data set. Results from both analyses consistently demonstrated that sample size has an insignificant effect on the coefficient values of reliability. Reliability increased as the number of items reached 18. Adding more items only negligibly increased the coefficients. Overall, the CA method consistently produced higher coefficient values of reliability compared to the S-B method.

Attitude to Health↗

Reliability of the National Institutes of Health Stroke Scale. Extension to non-neurologists in the context of a clinical trial.

BACKGROUND AND PURPOSE: The reliability of the National Institutes of Health Stroke Scale (NIHSS) has been established through testing its use in live and videotaped patients. This reliability testing has primarily focused on the use of the scale by neurologists. We sought to determine the reliability of the NIHSS as used by non-neurologists in the context of a clinical trial. METHODS: In anticipation of the initiation of a randomized trial of a new therapy for patients with acute ischemic stroke, 30 physician investigators (30% of whom were not neurologists) and 29 non-physician study coordinators were trained in the use of the NIHSS at an informational and training conference using standardized videotaped patient examinations. A series of 4 patients were rated initially. After 3 months, the same 4 patients were rerated, providing a measure of intraobserver reliability. An additional series of 4 new patients were also rated after 3 months and, with the initial 4 ratings, provided data for assessment of interobserver reliability. RESULTS: Overall, 28% of the raters had previous experience with the NIHSS, and 22% had previously used the videotapes as used in the present trial. The coefficients of determination (r2) were each greater than .95 when the means of the two ratings of the same 4 cases were compared between (1) neurologists and other types of physicians, (2) physicians and study coordinators, (3) raters who had prior experience with the NIHSS and those without prior experience, and (4) raters who had used the videotapes in the past and those who had never viewed the tapes. The calculated r2s were greater than .98 for the initial rating of the first 4 cases and for the later rating of the 4 new cases. The slopes of the regression lines were all near 1, indicating that the raters were similarly calibrated. The intraclass correlation coefficients were .93 and .95, reflecting high levels of intraobserver and interobserver reliability. CONCLUSIONS: These data extend the previously demonstrated reliability of the NIHSS to non-neurologists and show that both a variety of physician investigators and nurse study coordinators can be rapidly trained to reliably apply the scale in the context of an actual clinical trial.

Cerebrovascular Disorders↗

Reliability of visual analog and verbal descriptor scales for "objective" measurement of temporomandibular disorder pain.

Eight dentists viewed standardized videotapes showing palpations of the temporomandibular joint and muscles of mastication and recorded their judgments concerning the amount of pain the patient was experiencing. Judgments were recorded using a four-point verbal descriptor scale (VDS) ("none", "mild", "moderate", "severe" pain) or a 100-mm visual analog scale (VAS) anchored with the terms "no pain" and "worst pain possible". Test/re-test reliability over a one-week period and interjudge reliabilities were calculated for each scale; reliabilities of the two scales were directly compared based on the statistical equivalence of weighted kappa and the Intraclass Correlation Coefficient. Neither scale showed satisfactory reliability. Median test/re-test reliabilities were k = 0.590 for the VDS and r = 0.822 for the VAS. Interjudge reliabilities averaged k = 0.394 for the VDS and r = 0.735 for the VAS. Direct comparison of reliabilities for the two scales showed no clear advantage for either scale. The marginal reliabilities of these scales, when used by dentists to quantify the patient's pain, suggest that neither scale should be regarded as an "objective" pain measure.

Decision Making↗

The reliability of the Diabetes Care Profile for African Americans.

The Diabetes Care Profile (DCP) is an instrument used to assess social and psychological factors related to diabetes and its treatment. The reliability of the DCP was established in populations consisting primarily of Caucasians with type 2 diabetes. This study tests whether the DCP is a reliable instrument for African Americans with type 2 diabetes. Both African American (n = 511) and Caucasian (n = 235) patients with type 2 diabetes were recruited at six sites located in the metropolitan Detroit area. Scale reliability was calculated by Cronbach's coefficient alpha. The scale reliabilities ranged from .70 to .97 for African Americans. These reliabilities were similar to those of Caucasians, whose scale reliabilities ranged from .68 to .96. The Feldt test was used to determine differences between the reliabilities of the two patient populations. No significant differences were found. The DCP is a reliable survey instrument for African American and Caucasian patients with type 2 diabetes.

Black or African American↗

Reliability of a measure of post-stroke shoulder pain in patients with and without aphasia and/or unilateral spatial neglect.

OBJECTIVE: To determine the inter/intra-rater reliability of expert physiotherapists (PTs) measuring post-stroke shoulder pain with 100 mm vertical visual analogue scales (VAS; intensity, frequency and affective response) and a categorical site-of-pain scale. DESIGN: Three PTs independently rated subjects (normal clinical procedure but with a standardized starting position) on three days, at the same time of day, during one week in a randomized order determined by a nested latin square. Reliability for VAS scores was determined with the intraclass correlation coefficient (ICC) and for site-of-pain with the kappa statistic (kappa). Acceptable reliability was set at 0.75. The limits of agreement were also calculated. SETTING: Community. SUBJECTS: Thirty-three patients, mean time post stroke 42 months (range 7-360). RESULTS: Mean inter-rater reliability was 0.79 for intensity, 0.75 for frequency and 0.62 for affective response (ICC). The limits of agreement were wide and rater bias was significant for 6/27 ratings. Mean intra-rater reliability was 0.70 for intensity, 0.77 for frequency and 0.69 for affective response (ICC). For site-of-pain inter-rater reliability ranged from 0.156 (kappa) to 0.385 (kappa) and intrarater reliability ranged from 0.300 (kappa) to 0.559 (kappa). CONCLUSIONS: Although inter-rater reliability was acceptable for intensity and frequency there was a consistently large systematic bias between pairs of raters. Agreement might be improved if a standardized assessment procedure was used and/or if training in pain behaviour interpretation was provided.

Aged↗

Reliability of assessment tools in rehabilitation: an illustration of appropriate statistical analyses.

OBJECTIVE: To provide a practical guide to appropriate statistical analysis of a reliability study using real-time ultrasound for measuring muscle size as an example. DESIGN: Inter-rater and intra-rater (between-scans and between-days) reliability. SUBJECTS: Ten normal subjects (five male) aged 22-58 years. METHOD: The cross-sectional area (CSA) of the anterior tibial muscle group was measured using real-time ultrasonography. MAIN OUTCOME MEASURES: Intraclass correlation coefficients (ICCs) and the 95% confidence interval (CI) for the ICCs, and Bland and Altman method for assessing agreement, which includes calculation of the mean difference between measures (d), the 95% CI for d, the standard deviation of the differences (SDdiff), the 95% limits of agreement and a reliability coefficient. RESULTS: Inter-rater reliability was high, ICC (3,1) was 0.92 with a 95% CI of 0.72 --> 0.98. There was reasonable agreement between measures on the Bland and Altman test, as d was -0.63 cm2, the 95% CI for d was -1.4 --> 0.14 cm2, the SDdiff was 1.08 cm2, the 95% limits of agreement -2.73 --> 1.53 cm2 and the reliability coefficient was 2.4. Between-scans repeatability was high, ICCs (1,1) were 0.94 and 0.93 with 95% CIs of 0.8 --> 0.99 and 0.75 --> 0.98, for days 1 and 2 respectively. Measures showed good agreement on the Bland and Altman test: d for day 1 was 0.15 cm2 and for day 2 it was -0.32 cm2, the 95% CIs for d were -0.51 --> 0.81 cm2 for day 1 and -0.98 --> 0.34 cm2 for day 2; SDdiff was 0.93 cm2 for both days, the 95% imits of agreement were -1.71 --> 2.01 cm2 for day 1 and -2.18 --> 1.54 cm2 for day 2; the reliability coefficient was 1.80 for day 1 and 1.88 for day 2. The between-days ICC (1,2) was 0.92 and the 95% CI 0.69 --> 0.98. The d was -0.98 cm2, the SDdiff was 1.25 cm2 with 95% limits of agreement of -3.48 --> 1.52 cm2 and the reliability coefficient 2.8. The 95% CI for d (-1.88 --> -0.08 cm2) and the distribution graph showed a bias towards a larger measurement on day 2. CONCLUSIONS: The ICC and Bland and Altman tests are appropriate for analysis of reliability studies of similar design to that described, but neither test alone provides sufficient information and it is recommended that both are used.

Adult↗

The effect of training on rater reliability on the scoring of the NART.

OBJECTIVES: This study investigates whether the accuracy of judging National Adult Reading Test (NART) words known to have lower inter-rater reliability can be improved by training and use of the pronunciation guide. DESIGN: Two groups (Experimental and Control), were compared with three repeated measures: Occasion (first and second i.e. 'post-training'), Word Reliability (high and low) and Pronunciation Guide (without and with guide). METHODS: Ten words were selected from the NART: five lower reliability and five high reliability words. These were presented aurally in correct and incorrect form to participants (N = 20) who judged correctness of pronunciation without or with a pronunciation guide. Each group repeated the task again, the Experimental group having received training. RESULTS: Accuracy was significantly worse for the low reliability words. The experimental group's accuracy was significantly better after training than the control group's and their own performance prior to training. The use of the guide enhanced accuracy, particularly for the low reliability words. CONCLUSION: Training in administration of the NART improves raters' accuracy and use of the pronunciation guide. This offers an alternative to the suggestion of improving the NART's reliability by replacing lower reliability words and therefore would avoid the need to re-standardize a modified test.

Adult↗

The reliability of Form 90: an instrument for assessing alcohol treatment outcome.

OBJECTIVE: Project MATCH is a randomized clinical trial consisting of five outpatient and five aftercare units at nine sites. Of importance in this multisite trial examining the efficacy of client-treatment matching was the cross- and within-site reliability of the structured interview used to assess alcohol treatment outcomes, the Form 90. Evaluation of the reliability of Form 90 is the subject of this article. METHOD: The reliability of Form 90 was evaluated in two test-retest studies. The cross-site reliability study consisted of 70 paired test-retest interviews conducted by different interviewers. Clients for this study were recruited from inpatient, outpatient and college settings. The within-site reliability study had a total of 108 paired test-retest interviews, with 54 of the retests conducted by different interviewers and 54 by the same interviewer. Clients for this study were most often presenting for alcohol treatment at the nine sites and were selected to be representative of the larger Project MATCH sample. RESULTS: Good-to-excellent reliability was found for all key summary measures of alcohol consumption and psychosocial functioning, and most frequently used illicit drugs had moderate reliability. No decay in consistency of self-reported drinking was found at more distal points from dates of test-retest interviews. Application of 68% confidence intervals for primary alcohol consumption measures suggests that trained researchers and clinicians can obtain consistent information regarding client drinking. CONCLUSIONS: Form 90 appears to be a reliable instrument for alcohol treatment assessment research when interviewers have received careful training and supervision in its use.

Adult↗

Testing reliability of plaque and gingival indices. Two methods.

This investigation was undertaken to compare two methods of interexaminer and intraexaminer reliability in the evaluation of Plaque and Gingival Indices prior to a study of toothbrushing. Inter-/intraexaminer reliabilities were compared using a projected slide series consisting of 40 slides of clinical examples of gingival inflammation and plaque accumulation. Time between assessments was three weeks. Using the slide technique, intraexaminer reliability was established for: (1) Gingival Indices and (2) Plaque Indices. Interexaminer reliability was also established for Gingival Indices. Interexaminer reliability could not be established for Plaque Indices on the first assessment but was established on the post-assessment. Intraexaminer reliability was also determined through clinical examinations of patients. A third clinician was used to manipulate the tissue while investigators evaluated bleeding on provocation and plaque accumulation. Significant results were established for the Gingival Indices and Plaque Indices. Results of this investigation suggest that significant inter-/intraexaminer reliabilities may be obtained for gingival indices using the slide technique. In addition, the clinic technique appeared useful for assessing interexaminer reliability for Gingival Indices. Plaque Indices using the slide technique required more practice than those using the clinic technique.

Dental Health Surveys↗

Measures of reliability in sports medicine and science.

Reliability refers to the reproducibility of values of a test, assay or other measurement in repeated trials on the same individuals. Better reliability implies better precision of single measurements and better tracking of changes in measurements in research or practical settings. The main measures of reliability are within-subject random variation, systematic change in the mean, and retest correlation. A simple, adaptable form of within-subject variation is the typical (standard) error of measurement: the standard deviation of an individual's repeated measurements. For many measurements in sports medicine and science, the typical error is best expressed as a coefficient of variation (percentage of the mean). A biased, more limited form of within-subject variation is the limits of agreement: the 95% likely range of change of an individual's measurements between 2 trials. Systematic changes in the mean of a measure between consecutive trials represent such effects as learning, motivation or fatigue; these changes need to be eliminated from estimates of within-subject variation. Retest correlation is difficult to interpret, mainly because its value is sensitive to the heterogeneity of the sample of participants. Uses of reliability include decision-making when monitoring individuals, comparison of tests or equipment, estimation of sample size in experiments and estimation of the magnitude of individual differences in the response to a treatment. Reasonable precision for estimates of reliability requires approximately 50 study participants and at least 3 trials. Studies aimed at assessing variation in reliability between tests or equipment require complex designs and analyses that researchers seldom perform correctly. A wider understanding of reliability and adoption of the typical error as the standard measure of reliability would improve the assessment of tests and equipment in our disciplines.

Algorithms↗

Intramachine and intermachine reliability for selected dynamic muscle performance tests.

The Cybex 6000 isokinetic dynamometer is a new isokinetic device for which no published reports of reliability have been presented in the literature. In addition, the manufacturer not only claims that the new Cybex 6000 is reliable but that torque data obtained from the Cybex 6000 are consistent with data obtained from past Cybex systems, such as the Cybex II. The purpose of this study was to investigate the intramachine reliability of the Cybex 6000 to itself and the intermachine reliability of the Cybex 6000 and the Cybex II. Data on peak torque, work, and power were collected using the Cybex 6000, and data on peak torque were obtained using the Cybex II for knee flexion and extension in 20 volunteers (10 males, 10 females). Subjects were tested three times, twice on the Cybex 6000 and once on the Cybex II, approximately 1 week apart across a 3-week period of time at angular velocities of 60, 180, and 300 degrees/sec. Data were analyzed using intraclass correlations. Results indicated that the majority of test-retest correlation coefficients for all parameters for intramachine reliability of the Cybex 6000 were above .90. Comparing peak torque obtained with the Cybex 6000 to that obtained with the Cybex II (intermachine reliability), correlation coefficients ranged from .72 to .89. In conclusion, information obtained on the Cybex 6000 appears to be quite reliable in a test-retest situation using the same equipment and moderately reliable when compared to the Cybex II. Clinical implications for these results are discussed.

Adult↗

Relationship of the pelvic angle to the sacral angle: measurement of clinical reliability and validity.

There is a need to better document the reliability and validity of assessment measures used in physical therapy. Studies documenting the reliability of measurement of the pelvic angle and its relationship to sacral motion are presently inconclusive. The purpose of this study was twofold. First, we wanted to determine the reliability and validity of a goniometric measurement of the pelvic angle. We also wanted to test the hypothesis that there is a relationship between the pelvic angle and the sacral angle. Intertester and intratester reliability of goniometric pelvic angle measurements of 23 healthy young adults were examined using three different raters. Radiographic measurements of the pelvic and sacral angle using two raters and goniometric measurement of the pelvic angle using a single rater were taken from 15 patients with low back pain who had been referred for X-rays. Intraclass correlation coefficients (ICCs) of intratester reliability for goniometric measurements of the pelvic angle were .93, .96, and .96. The intertester reliability was .95. The ICCs for intratester reliability for radiological measurements were .92 and .95 for the sacral angle and .98 for both measurements of the pelvic angle. Intertester reliability coefficients were .86 and .88, respectively. The Pearson correlation coefficients for the goniometric and radiological measurements of the pelvic angle were .85 and .68. A comparison of the radiological and goniometric measurements of the pelvic angle with the sacral angle demonstrated low average correlations of .43 and .58, respectively. The results indicate a high level of correlation between and within testers for goniometric measurements of the pelvic angle but only a fair correlation between goniometric and radiological measurements of the pelvic angle.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult↗

Reliability of McConnell's classification of patellar orientation in symptomatic and asymptomatic subjects.

STUDY DESIGN: Test-retest reliability study with blinded testers. OBJECTIVES: To determine the intratester reliability of the McConnell classification system and to determine whether the intertester reliability of this system would be improved by one-on-one training of the testers, increasing the variability and numbers of subjects, blinding the testers to the absence or presence of patellofemoral pain syndrome, and adhering to the McConnell classification system as it is taught in the "McConnell Patellofemoral Treatment Plan" continuing education course. BACKGROUND: The McConnell classification system is currently used by physical therapy clinicians to quantify static patellar orientation. The measurements generated from this system purportedly guide the therapist in the application of patellofemoral tape and in assessment of the efficacy of treatment interventions on changing patellar orientation. METHODS AND MEASURES: Fifty-six subjects (age range, 21-65 years) provided a total of 101 knees for assessment. Seventy-six knees did not produce symptoms. A researcher who did not participate in the measuring process determined that 17 subjects had patellofemoral pain syndrome in 25 knees. Two testers concurrently measured static patellar orientation (anterior/posterior and medial/lateral tilt, medial/lateral glide, and patellar rotation) on subjects, using the McConnell classification system. Repeat measures were performed 3-7 days later. A kappa (kappa) statistic was used to assess the degree of agreement within each tester and between testers. RESULTS: The kappa coefficients for intratester reliability varied from -0.06 to 0.35. Intertester reliability ranged from -0.03 to 0.19. CONCLUSION: The McConnell classification system, in its current form, does not appear to be very reliable. Intratester reliability ranged from poor to fair, and intertester reliability was poor to slight. This system should not be used as a measurement tool or as a basis for treatment decisions.

Adult↗

Reliability of some tremor measurement outcome variables in field testing situations.

OBJECTIVE: Many summary measures of data obtained from tremor measurement procedures are commonly reported. The reliability of many of these summary measures of tremor measurements made in field testing situations is unknown. The purpose of the present investigation was to assess the reliability of a number of summary measures produced by the software of a widely used, commercially available tremor measurement instrument using data collected in three field epidemiologic studies. METHODS: Tremor data were obtained from 689 participants in 3 previously conducted studies of groups exposed to elemental mercury or arsenic. A widely used, commercially available tremor measurement instrument was used. Two-axis accelerometer measurements were obtained on 2 or more trials from each hand for each participant. Estimates of trial-to-trial and internal consistency reliability were calculated for 5 summary measures calculated by instrument manufacturer's software and 5 additional summary measures calculated from data output by the software. RESULTS: An RMS acceleration measure had the highest reliability in all 3 studies. The average over 4 trials of RMS acceleration and its logarithm had high reliability (>0.9). Recalculation of a tremor summary index and a harmonicity index as suggested by Edwards and Beuter (1999) resulted in measures with higher reliability and better distributional shape than the corresponding measures provided by the instrument manufacturer's software. The results in all three studies were similar. CONCLUSIONS: For the tremor measurement instrument and testing procedure that we employed, we recommend using the common logarithm of the RMS accelerations and recalculated tremor index as summary measures. We also recommend employing multiple trials of each type (e.g., with each hand) and averaging summary measures from those trials to derive outcome measures of tremor for use in epidemiologic studies. We recommend at least 2 trials for RMS acceleration measures and more for less reliable measures, particularly for designs employing repeated measurements of individuals. Summary measures averaged over at least 4 trials for mean frequency, dispersion of frequency, and power in the 3-6.5 and 6.6-10 Hz frequency ranges have sufficiently high reliability for use in epidemiologic studies.

Aged↗