Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

The reliability of observational data: I. Theories and methods for speech-language pathology.

Much research and clinical work in speech-language pathology depends on the validity and reliability of data gathered through the direct observation of human behavior. This paper reviews several definitions of reliability, concluding that behavior observation data are reliable if they, and the experimental conclusions drawn from them, are not affected by differences among observers or by other variations in the recording context. The theoretical bases of several methods commonly used to estimate reliability for observational data are reviewed, with examples of the use of these methods drawn from a recent volume of the Journal of Speech and Hearing Research (35, 1992). Although most recent research publications in speech-language pathology have addressed the issue of reliability for their observational data to some extent, most reliability estimates do not clearly establish that the data or the experimental conclusions were replicable or unaffected by differences among observers. Suggestions are provided for improving the usefulness of the reliability estimates published in speech-language pathology research.

Humans↗

The reliability of the wolf motor function test for assessing upper extremity function after stroke.

OBJECTIVE: To examine the reliability of the Wolf Motor Function Test (WMFT) for assessing upper extremity motor function in adults with hemiplegia. DESIGN: Interrater and test-retest reliability. SETTING: A clinical research laboratory at a university medical center. PATIENTS: A sample of convenience of 24 subjects with chronic hemiplegia (onset >1yr), showing moderate motor impairment. INTERVENTION: The WMFT includes 15 functional tasks. Performances were timed and rated by using a 6-point functional ability scale. The WMFT was administered to subjects twice with a 2-week interval between administrations. All test sessions were videotaped for scoring at a later time by blinded and trained experienced therapists. MAIN OUTCOME MEASURE: Interrater reliability was examined by using intraclass correlation coefficients and internal consistency by using Cronbach's alpha. RESULTS: Interrater reliability was.97 or greater for performance time and.88 or greater for functional ability. Internal consistency for test 1 was.92 for performance time and.92 for functional ability; for test 2, it was.86 for performance time and.92 for functional ability. Test-retest reliability was.90 for performance time and.95 for functional ability. Absolute scores for subjects were stable over the 2 test administrations. CONCLUSION: The WMFT is an instrument with high interrater reliability, internal consistency, test-retest reliability, and adequate stability.

Adolescent↗

Reliability of cervical range of motion using the OSI CA 6000 spine motion analyser on asymptomatic and symptomatic subjects.

Cervical range of motion (ROM) is evaluated in both clinical and research settings. This study's purpose was to determine if ROM data obtained with the OSI CA 6000 Spine Motion Analyser (SMA) from asymptomatic and symptomatic cervical subjects were reliable within and between testers. Cervical ROM was measured in all three planes in 30 adult asymptomatic and 20 adult symptomatic subjects. A standardized protocol was used to fit each subject with the OSI SMA cervical hardware. Subjects were tested in a seated position with the trunk stabilized. Subjects performed four trials of each pain-free cervical motion during testing. The hardware was completely removed and replaced by the same tester and ROM trials in all three planes were repeated for intratester asymptomatic and symptomatic reliability. The same procedure was completed by a second tester for asymptomatic intratester and intertester reliability. Repeated measures analysis of variance and intraclass correlation coefficients (ICC [2,1 and 2 k]) were used to analyse intra- and intertester reliability data. Intratester ICCs were 0.85 or higher (except for flexion 0.76) for asymptomatic subjects and 0. 87 or higher (except for flexion 0.68) for symptomatic subjects for all motions. Intertester ICCs were 0.88 or higher for all motions. Standard error of measurements were less than 3.92 degrees for all motions. Measures of cervical spinal ROM obtained with the OSI SMA showed good intertester reliablity for all motions, and good intratester reliability for all motions with the exception of the motion of flexion for one of the examiners, which showed moderate reliability.

Adult↗

[Determination of Reliability of Psychometric Tests in Psychiatry Using Canonical Correlation]

Test results (raw scores) are composed of an unknown true score and an error term. The error term can be estimated by means of test reliability which is defined by the ratio of true variance and obtained variance. Different estimates of reliability either based on single measurements (e. g. Cronbach's coefficient, split half reliability, Kuder Richardson method) or two measurements (test/retest, inter- or intrarater reliability) are available. Parallel test reliability depends on the correlation of two different tests obtained in one session. Canonical correlation methods allow an extension of the parallel test situation and split half technique. Two or more tests are performed in a sample of subjects. Randomized subsets are correlated using canonical correlation technique. The objective of this study is to estimate the homogeneity of test batteries. 94 patients (64 f, 30 m; age: 54 - 89 ys.) supposed to have dementia were tested using the clocktest (CT, scores: 1 - 5), MMSE (mini mental state examination) and SKT (Syndrom Kurztest). Four (i, j: 1 - 4) subsets of 20 patients each were determined by random and the following characteristics were calculated: Empiric correlation coefficient for n = 94 (R), canonical correlation coefficient (Rcan), eigenvalues (EV) and redundancy (Rnd) of corresponding variable sets. The results of canonical analysis showed canonical correlation coefficients in order of 0.8 to 0.9 (p-values < 0,001). This high internal consistency can be interpreted as a measure of reliability of the test batteries. In conclusion, canonical correlation based on parallel tests splitted in subsets gives information on consistency, i. e. reliability, of test batteries in addition to conventional correlation methods.

Journal Article↗

[Reliability of quantifying vascular white matter brain lesions -- a contribution to reproducible quantitative diagnosis].

PURPOSE: Microangiopathic lesions of the brain tissue correlate with the clinical diagnosis of vascular subcortical dementia. The "experience-based" evaluation is insufficient. Rating scales may contribute to reproducible quantification. MATERIALS AND METHODS: In MRI studies of 10 patients, 9 neuroradiologists quantified vascular white matter lesions (WMLs) at two different points in time for 12 anatomically defined regions with respect to number, size and localization (score). For 9 observers and 10 studies, 90 intra-observer differences were obtained for each of the 12 WML scores. To calculate the inter-observer reliability, rating pairs were formed. Furthermore, 360 differences were computed for each score and rating for 12 anatomically defined WML scores, and the intraclass correlation (ICC) was calculated as a measure of agreement (reliability). RESULTS: As to the intra-observer reliability, the median of the differences was 1.5 for the entire brain as opposed to 0 for defined brain regions. The corresponding values for the inter-observer reliability were 3 and 1, respectively. The mean intra-class correlation coefficient for the 10 studies was 0.88, whereas the mean interclass correlation concerning the inter-observer reliability was 0.70, with the first and second rating being averaged. The rating of each study took about 6 minutes. CONCLUSION: The rating scale with high intra- and inter-observer reliability can dependably quantify WMLs and correlates with the clinical diagnosis of vascular dementia. Using a reliable rating scale, the diagnostic distinction of age-associated physiological vs. pathological size of the WML can make a contribution to the reproducible quantifiable diagnostic evaluation of vascular brain tissue lesions within the framework of dementia diagnostics.

Aged↗

Reliability and within subject variability of VE, VO2, heart rate and blood pressure during submaximum cycle ergometry.

To assess the reliability and within subject variability of steady-rate ventilation (VE), oxygen uptake (VO2), heart rate, systolic and diastolic blood pressure, 4 subjects exercised for 10 minutes at 3 work rates on a bicycle ergometer: 50 W, 125 W and 55% of maximum work rate (55% max). Each testing session included two work rates and only 2 testing sessions were scheduled per week. The order of the work rates was counterbalanced. In 8 to 10 weeks, 3 of the subjects completed 20 trials at 50 W while the fourth subject completed 11 trials, and all the subjects completed 10 trials at 125 W and 55% max. The within subject variability (S2w) was expressed as a percent of the mean steady-rate response. VO2 ranged from 21.2% to 27.5% of VO2max at 50 W, from 37.7% to 49.7% at 125 W and from 42.9% to 63.7% at 55% max. The S2w averaged 6.8% for VE, 4.3% for VO2, 3.2% for heart rate, 7.3% for systolic blood pressure and 10.5% for diastolic blood pressure. Reliability coefficients were calculated for the steady-rate scores by dividing the between subject variation by the total variation. The reliability was similar for VE, VO2 and heart rate and ranged from r = 0.69 to r = 0.97. Systolic and diastolic blood pressure reliabilities were lower and ranged from r = 0.27 and r = 0.80. In summary, the steady-rate ventilation, oxygen uptake and heart rate responses were reliable and consistent. The reliability of blood pressure was low. It is possible that this low reliability may result from variability in stroke volume or total peripheral resistance.

Adult↗

Reliability of chiropractic methods commonly used to detect manipulable lesions in patients with chronic low-back pain.

OBJECTIVE: To assess the intraexaminer and interexaminer reliability of a multidimensional spinal diagnostic method commonly used by chiropractors. DESIGN: An intraexaminer and interexaminer Latin square, repeated measures reliability study. The techniques of diagnosis under investigation included visual postural analysis, pain description by the patient, plain static erect x-ray film of the lumbar spine, leg length discrepancy, neurologic tests, motion palpation, static palpation, and orthopedic tests. PARTICIPANTS: Three experienced chiropractors examined 19 patients, and 2 experienced chiropractors examined 10 and 9 patients, respectively, who were suffering from chronic mechanical low-back pain. RESULTS: Intraexaminer reliability of the decision to manipulate a certain spinal segmental level was moderate (kappa = 0.47). The interexaminer agreement pooled across all spinal joints indicated fair agreement (kappa = 0.27). Interexaminer reliability for individual examiner pairs for the L4/L5 segmental level was slight (kappa = 0.09). At the L5/S1 level, the interexaminer reliability was fair (kappa = 0.25). For the sacroiliac joints, interexaminer reliability was slight (kappa = 0.04 and 0.14). CONCLUSION: This study of commonly used chiropractic diagnostic methods in patients with chronic mechanical low-back pain to detect manipulable lesions in the lower thoracic spine, lumbar spine, and the sacroiliac joints has revealed that the measures are not reproducible. The implementation of these examination techniques alone should not be seen by practitioners to provide reliable information concerning where to direct a manipulative procedure in patients with chronic mechanical low-back pain.

Adult↗

The kinematics of motion palpation and its effect on the reliability for cervical spine rotation.

BACKGROUND: The reliability of a test depends on its standardization. Instrumental measurement of the reproducibility of the test is an effective way to evaluate the level of standardization obtained. Improved standardization is believed to yield greater reliability. OBJECTIVE: The objectives of this study were to measure the technical ability of an examiner to reproduce the kinematics of motion palpation for cervical spine rotation and to evaluate the effect of standardization on the reliability of the test. DESIGN: A study of reproducibility of the kinematics of the test for cervical spine rotation was conducted by means of a computerized system of analysis of movement. The reliability when reproducibility was achieved was compared with reliability when it failed. RESULTS: The data collected enable us to establish a standardized protocol for the execution of the test. The standardized palpation is executed within 6 of inclination from the pure plane of rotation. The successful reproduction of the kinematics of the test raises its reliability to detect the presence of fixations (kappa raising from 0.337 and 0.352 to 0.682). CONCLUSIONS: A greater reliability, arising from a high level of reproducibility, enables us to document the advantages of the standardization of motion palpation in chiropractic.

Adult↗

A reliability study of clinical tooth wear measurements.

STATEMENT OF PROBLEM: Most studies examining tooth wear severity have been performed on dental casts. This indirect approach has limited applicability to dental practice because during the assessment of the casts, the identification of dentin exposure is difficult or even impossible. PURPOSE OF STUDY: The purpose of this study was to assess occlusal and incisal tooth wear clinically to determine the reliability of the assessment procedure and to establish the influence of selected relevant clinical variables (dental quadrant, tooth type, and severity of wear) on the reliability. MATERIAL AND METHODS: Forty-five volunteers (17 men, 28 women; mean age 33.7 +/- 10.7 years), 32 with temporomandibular disorders and 13 free from signs and symptoms of such disorders, were evaluated on 4 occasions. Two trained observers graded tooth wear at 2 different points in time with a 5-point ordinal scale developed for use in this study. The inter-rater and intra-rater reliability of the scale was expressed as Cohen's kappa. The influence of 2 clinical variables, dental quadrant and tooth type, on the values of kappa was tested with 1-way analysis of variance and post hoc Bonferroni tests. Probability levels of P< .05 were considered statistically significant. The influence of the final clinical variable, severity of wear, was assessed qualitatively. RESULTS: The overall values of the inter-rater and intra-rater reliability were substantial (kappa = 0.632 to 0.678). The clinical variable dental quadrant did not influence the kappa values, whereas the inter-rater reliability during the first session was better for incisors and canines than for premolars (1-way analysis of variance: F(3,23)=4.577, P=.012; post hoc Bonferroni tests: P=.030 and.036). Qualitative assessment of severity of wear indicated that the more advanced the tooth wear, the more reliably it could be graded. CONCLUSION: By means of the developed 5-point ordinal scale and within the limitations of this study, it was concluded that tooth wear can be assessed reliably in the clinical dental setting.

Adult↗

Topology of biological networks and reliability of information processing.

Survival of living cells and organisms is largely based on highly reliable function of their regulatory networks. However, the elements of biological networks, e.g., regulatory genes in genetic networks or neurons in the nervous system, are far from being reliable dynamical elements. How can networks of unreliable elements perform reliably? We here address this question in networks of autonomous noisy elements with fluctuating timing and study the conditions for an overall system behavior being reproducible in the presence of such noise. We find a clear distinction between reliable and unreliable dynamical attractors. In the reliable case, synchrony is sustained in the network, whereas in the unreliable scenario, fluctuating timing of single elements can gradually desynchronize the system, leading to nonreproducible behavior. The likelihood of reliable dynamical attractors strongly depends on the underlying topology of a network. Comparing with the observed architectures of gene regulation networks, we find that those 3-node subgraphs that allow for reliable dynamics are also those that are more abundant in nature, suggesting that specific topologies of regulatory networks may provide a selective advantage in evolution through their resistance against noise.

Biological Evolution↗

Recalibration improves inter-examiner reliability of TMD examination.

OBJECTIVE: The purpose of this study was to assess whether recalibration of examiners would improve the reliability of gathering clinical findings and related diagnoses of temporomandibular disorders (TMD) in accordance with the Research Diagnostic Criteria for TMD (RDC/TMD). MATERIAL AND METHODS: Two clinicians independently examined a total of 48 symptomatic and asymptomatic subjects according to the RDC/TMD on two occasions: examination 1 (E1). Aarhus, Denmark (n=24; 18 female, ages 18-59 years); examination 2 (E2). Malmö, Sweden (n=24; 18 female, ages 18-86 years). The clinicians were calibrated in the use of the RDC/TMD Axis-I examination on the day before E1. Six months later, they were recalibrated on the day before E2. Intra-class correlation coefficients (ICCs) were used to examine the inter-examiner reliability of the two clinicians on the two occasions (E1, E2). RESULTS: The intra-class correlation coefficients of vertical range of jaw motion differed little between E1 and E2. At E2, all other examination components consistently improved in reliability relative to E1. Similar improvements were seen for the frequently occurring RDC/TMD clinical diagnoses: Ia. Myofascial pain [ICC = 0.83 (E1) and 1.00 (E2)], IIa. Disk displacement with reduction [ICC = 0.26 (E1) and 0.64 (E2)], and IIIa. Arthralgia [ICC = 0.16 (E1) and 0.73 (E2)]. CONCLUSION: Recalibration considerably improved inter-examiner reliability for assessing RDC/TMD clinical variables and diagnoses, which are critically dependent on reliable assessment of clinical signs; improvement was most marked when initial inter-examiner reliability was low. Final inter-examiner reliabilities after recalibration were all associated with acceptable to excellent levels.

Adolescent↗

Reliability and motor memory.

The generalizability framework was employed to estimate variance components and reliability of performance in a positioning task. The subjects (n =70) performed 16 trials at each of three target lengths of 25 cm, 45 cm, and 65 cm. The data were analyzed using a subjects x targets x trials repeated measures ANOVA. Two reliability coefficients were estimated. The first (R) provided an estimate of performance reliability generalized over targets and trials. The second (R') treated targets as a source of true-score variance and hence was generalized over trials alone. Both reliability coefficients were higher for algebraic error than for absolute error, and R1 provided higher reliability estimates than R for both dependent variables. Increasing the number of positioning trials from 2 to 16 at each target length did not appreciably alter either reliability coefficient. The overall low reliability of R appears to compound the dependent variables and statistical power problems associated with short-term motor-memory studies.

Journal Article↗

Reliability and smallest detectable change determination for serratus anterior muscle strength and endurance tests.

The purposes of this study were to determine inter-tester reliability, one-week test-retest reliability and smallest detectable difference (SDD) of serratus anterior muscle strength and endurance tests. Asymptomatic subjects were tested on an apparatus designed by the investigators. During strength testing, subjects performed isometric contractions recorded by a hand-held dynamometer. For endurance testing, subjects held a dumbbell of 15% of their body weight and performed repetitions until they became fatigued. Intraclass correlation coefficients (ICC) for the strength test revealed good inter-tester reliability (ICC2,3 = .90-.93) and good one-week test-retest reliability (ICC2,3 = .83-.89). For the endurance test, ICCs showed good inter-tester reliability (ICC2,1 = .71-.76) but moderate one-week test-retest reliability (ICC2,1 = .59-.62). The SDDs at 68% confidence level ranged from 22.7 to 39.2 newtons for the strength test, and 11 to 20 repetitions for the endurance test. In summary, the technique used in the study is reliable for quantifying the SA muscle strength.

Electromyography↗

Reliability of speech and language therapists using therapy outcome measures.

The Therapy Outcome Measure (TOM) aims to provide Speech and Language Therapists (SLT) with a practical tool to measure outcomes of care by providing a quick and simple measure which can be used over time with patients and clients in a routine clinical setting. The TOM allows therapists to reflect their clinical judgement on the dimensions of impairment, disability/activity, handicap/participation and well-being on an 11-point ordinal scale. The purpose of this paper is to examine the reliability and the influences on reliability of SLT using this measure. Three studies are presented and give information on 73 SLT using the measure with different client groups. Study one assesses the degree of reliability of six SLT using the TOM following training and practice. Reliability was studied on three occasions and the results demonstrate the influence of training and practice. Study two included 56 SLT to examine reliability over a broader range of client groups and to investigate the effect of the SLT specializing on rating patients from within or -out that specialism. Eleven therapists were included in the study, which examined the influence of a specific training approach. The participating SLT achieved a substantial-moderate level of reliability; this was established in all domains. The degree of reliability achieved on the TOM was affected by some, but not extensive, training and experience.

Clinical Protocols↗

Reliability of the Prognos electrodermal device for measurements of electrical skin resistance at acupuncture points.

OBJECTIVES: (1) To characterize and calibrate an electrodermal screening device, Prognos. (2) To replicate a previous test-retest reliability study of this device with measurements of electrical skin resistance (ESR) at 24 Jing-well acupuncture points (APs). (3) To determine measurement precision in three successively more exacting trial protocols on the same set of subjects. SETTINGS: Oregon College of Oriental Medicine and Portland State University, Portland, OR. INSTRUMENTS: The Prognos device was electrically characterized by a team of research engineers at the Biomedical Signal Processing Laboratory of Portland State University. They determined that Prognos measures the average direct-current (DC) resistance between a metallic wrist strap and an electrode probe tip. The probe tip is connected to a linear spring set to trigger with an optically generated signal at a deflection of 2.62 mm, which corresponds to an average applied force of 2.68 +/- 0.04 N (mean +/- standard deviation [SD], n = 6). They also determined that the device quantifies resistance by applying a 1.1 microA current for an average of 223 +/- 3 ms (n = 7). When calibrated against a series of known resistors, Prognos measures accurately in the range of 150 kOmega to 14.3 MOmega with an error of less than 0.4%. SUBJECTS: Thirty-one (31) healthy volunteers, 17 females and 14 males, 23-63 years of age. RESULTS OF RELIABILITY TEST-RETEST: The mean reliability of a single measurement was; 0.758 for a standard measurement protocol of four sequential sweeps of 24 Jing-well (Ting) APs; 0.851 for four sequential sweeps after ink-marking the APs; and 0.961 for four rapid repeat measurements at each inked AP. Mean absolute values of ESR decreased between the standard and marked protocols, but not between the marked and rapid repeat protocols. CONCLUSIONS: Prognos performs accurately, against known resistors over the reported range of ESR. The reliability in the standard protocol (r = 0.758) is comparable to the reliability of 0.721 demonstrated under similar conditions by other investigators. Marking APs, and performing measurements in a rapid sequence, increases reliability of ESR measurements. Increased reliability in the second and third protocols is associated with decreased mean ESR values which may be related to increased accuracy of Prognos probe placement and/or inking the APs.

Acupuncture Points↗

Structured interview and uniform assessment improves diagnostic reliability.

OBJECTIVE: To compare a Childhood Uniform Assessment Package (CUAP), including a computerized structured diagnosis, with routine assessment and treatment in public mental health settings. DATA SOURCES/STUDY SETTINGS: Data was collected prospectively on 250 children and adolescents in both public mental health inpatient and outpatient settings in a large metropolitan area and a rural area. STUDY DESIGN: Subjects were randomized to either routine assessment and treatment as usual (ATU) or ATU plus an additional "gold standard" assessment battery Childhood Uniform Assessment Package (CUAP). Outcome measures were taken at admission (baseline), discharge, and again 6 months later. METHODS: The study was conducted at a State Hospital (CUAP, n = 75; ATU, n = 75) and a Community Mental Health center (CUAP, n = 50; ATU, n = 50). The "gold standard" diagnostic process was established at the Children's Medical Center-Dallas. Research focused on a comparison of the CUAP diagnostic process to the existing diagnostic process (ATU) and the service delivery system of an inpatient and outpatient public sector clinical treatment setting. PRINCIPAL FINDINGS: A bachelor's level individual can be trained to administer a highly reliable diagnostic battery to meet a "gold standard," suggesting a possible cost-effective way to assist in diagnostic evaluations. Higher reliability was found between this standardized assessment package (CUAP) and inpatient physicians than for outpatient physicians. The highest interrater reliabilities were found for attention deficit and substance abuse disorders, less so for the other behavior disorders. The use of CUAP results in more reliable diagnoses in public settings than those provided by typical clinical staff by identifying mood and anxiety disorders (disorders with the lowest reliability) with better reliability. The addition of "gold standard" diagnostic assessments (CUAP) did not appear to affect length of stay, number of medication changes, use of seclusion or restraints, and other behavioral interventions in the inpatient setting. Outpatient follow-up services did not differ for CUAP versus ATU either. CONCLUSIONS: A standard uniform assessment package that includes a structured diagnostic instrument can improve overall diagnostic reliability but may not have a significant overall impact in clinical treatment strategies or outcomes without additional intervention to assure proper use of the information. A well-trained bachelor's level assistant can administer such a battery.

Adolescent↗

Reliability of the Barthel Index when used with older people.

OBJECTIVE: the Barthel Index (BI) has been recommended for the functional assessment of older people but the reliability of the measure for this patient group is uncertain. To investigate this issue we undertook a systematic review to identify relevant studies from which an overview is presented. METHOD: studies investigating the reliability of the BI were obtained by searching Medline, Cinahl and Embase to January 2003. Screening for potentially relevant papers and data extraction of the studies meeting the inclusion criteria were carried out independently by two researchers. RESULTS: the scope of the 12 studies identified included all the common clinical settings relevant to older people. No study investigated test-retest reliability. Inter-rater reliability was reported as 'fair' to 'moderate' agreement for individual BI items, and a high percentage agreement for the total BI score. However, these findings were difficult to interpret as few studies reported the prevalence of the disability categories for the study populations. There may be considerable inter-observer disagreement (95% CI of +/-4 points). There was evidence that the BI might be less reliable in patients with cognitive impairment and when scores obtained by patient interview are compared with patient testing. The role of assessor training and/or guidelines on the reliability of the BI has not been investigated. CONCLUSIONS: although the BI is highly recommended, there remain important uncertainties concerning its reliability when used with older people. Further studies are justified to investigate this issue.

Activities of Daily Living↗

How valid and reliable are patient satisfaction data? An analysis of 195 studies.

OBJECTIVE: To assess the properties of validity and reliability of instruments used to assess satisfaction in a broad sample of health service user satisfaction studies, and to assess the level of awareness of these issues among study authors. DESIGN: Examination and analysis of 195 papers published in 1994 in 139 journals. The following databases were searched: British Nursing Index, CINAHL, EMBASE, MedLine, Popline, and PsycLIT. MAIN MEASURES: Number and types of strategies used for content, criterion, and construct validity, and for stability and internal consistency. Associations between validity/reliability and other study characteristics. RESULTS: Eighty-nine (46%) of the 195 studies reported some validity or reliability data; 76 reported some element of content validity; 14 reported criterion validity, with patient's intent to return the most commonly used criterion; four reported construct validity. Thirty-four studies reported internal consistency reliability, 31 of which used Cronbach's coefficient alpha; eight studies reported test-retest reliability. Only 11 studies (6% of the 181 quantitative studies) reported content validity and criterion or construct validity and reliability. 'New' instruments designed specifically for the reported study demonstrated significantly less evidence for reliability/validity than did 'old' instruments. CONCLUSION: With few exceptions, the study instruments in this sample demonstrated little evidence of reliability or validity. Moreover, study authors exhibited a poor understanding of the importance of these properties in the assessment of satisfaction. Researchers must be aware that this is poor research practice, and that lack of a reliable and valid assessment instrument casts doubt on the credibility of satisfaction findings.

Data Collection↗