Search PubMedSearch

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Reliability of the Modified Ashworth Scale in the assessment of plantarflexor muscle spasticity in patients with traumatic brain injury.

Although the Modified Ashworth Scale (MAS) is commonly used to assess the severity of muscle spasticity for ankle plantarflexors, its reliability has only been established for elbow muscles. Interrater reliability, intrarater reliability and temporal (between-days) reliability were examined in this study. Also, interrater reliability for use of the scale with plantarflexors was compared with reported results from the measurement of elbow flexors. Thirty adult volunteers with traumatic brain injuries participated. There were 20 men and 10 women; the mean age was 28.3 years (SD = 10.8). Two physical therapists used the MAS to score the subjects independently. Measurements were repeated to yield multiple scores for intrarater reliability assessment. Twenty-one of the subjects returned individually on separate days to be measured again, so that temporal reliability could be assessed. Spearman's correlation coefficients were 0.73 for interrater reliability 0.74 and 0.55 for intrarater reliability, and 0.82 for temporal reliability. Overall, reliability of the MAS for assessing plantarflexor spasticity in patients with traumatic brain injury was found to be minimally adequate to support its continued use. However, interrater reliability was less than that which has been reported for elbow flexors, and intrarater reliability findings were mixed.

Adolescent

Reliability and validity of the chronic respiratory questionnaire (CRQ).

BACKGROUND: The Chronic Respiratory Questionnaire (CRQ) is frequently applied to assess quality of life in patients with chronic obstructive pulmonary disease (COPD). However, the reliability and validity of this questionnaire have not yet been determined. This study investigates the reliability and validity of the four separate dimensions of the CRQ. METHODS: The CRQ was administered on two consecutive days to 40 patients with COPD (mean FEV1 44% predicted, FEV1/IVC 37% predicted). Internal consistency reliability of each dimension was investigated by Cronbach's alpha reliability coefficient, test retest reliability by the Spearman-Brown reliability coefficient (p), and content validity by Pearson's correlation coefficient between the CRQ and the symptom checklist (SCL-90). RESULTS: Items of the fatigue, emotion, and mastery dimensions showed a high internal consistency reliability (alpha = 0.71-0.88) as well as a high test retest reliability (p above 0.90). These three dimensions correlated with comparable dimensions of the SCL-90. Items of the dyspnoea dimension showed a low internal consistency reliability (alpha = 0.53) and a test retest reliability of p = 0.73. CONCLUSIONS: Items of the dimensions fatigue, emotion, and mastery of the CRQ are reliable and valid and can be used to assess quality of life in patients with severe airways obstruction. Items of the dyspnoea dimension are less reliable and should not be included in the overall score of the CRQ in comparative research. However, by scoring the items of dyspnoea separately they may be useful for the evaluation of the effects of intervention in a specific patient.

Aged

A study of the test-retest reliability of ten olfactory tests.

Ten tests of olfactory function (including tests of odor identification, detection, discrimination, memory, and suprathreshold odor intensity and pleasantness perception) were administered on two test occasions to 57 subjects ranging in age from 18 to 83 years. The stability of the average test scores was determined across the two test sessions for 14 measures derived from these 10 tests and for subcomponents of the Japanese T&T olfactometer threshold test. In addition, the test-retest reliability (Pearson r) of each test measure was established. With the exception of a response bias measure, the average test scores did not differ significantly across the two test sessions. Statistically, the reliability coefficients of the primary test measures fell into three general classes bound by the following r values: 0.43-0.53; 0.67-0.71; 0.76-0.90. Detection threshold values were more reliable than recognition threshold values; those based upon a single ascending presentation series were much less reliable than those based upon a staircase procedure. The relationship between test length and reliability was examined for several of the tests and mathematically modeled. For example, within the staircase series incorporating the odorant phenyl ethyl alcohol, reliability was related (R2 = 0.984) to the number of reversals included in the threshold estimate by a function derived from the Spearman-Brown formula; namely, reliability = 0.455* # reversals/[1 + 0.455 (# reversals - 1)]. Reversal location, per se, had little influence on reliability. Overall, this study suggests that (i) considerable variation is present in the reliability of olfactory tests, (ii) reliability is a function of test length, and (iii) caution is warranted in comparing results from nominally different olfactory tests in applied settings since the findings may, in some instances, simply reflect the differential reliability of the tests.

Adolescent

Interrater reliability of auscultation of breath sounds among physical therapists.

BACKGROUND AND PURPOSE: Although auscultation is routinely used in the assessment of respiratory status, the ability of the rater to accurately and consistently identify lung sounds has been questioned. The literature on this issue is sparse and has focused on reliability of auscultation of tape-recorded rather than in vivo lung sounds. The purposes of this study were to determine the interrater reliability of physical therapists in the direct auscultation of lung sounds based on their clinical experience in chest physical therapy and to determine whether the adoption of standardized nomenclature and education on proper technique and interpretation affects reliability. SUBJECTS AND METHODS: A group of 57 registered physical therapists were stratified by clinical experience into four groups. Sixteen therapists (ie, 4 in each stratum) were randomly chosen using a random number table. The following criteria were developed to delineate clinical experience. Group 1 subjects were senior chest physical therapists with at least 5 years of experience in this area of practice. Group 2 subjects were experienced therapists who had a minimum of 2 years of experience in chest physical therapy and were currently practicing in this area. Group 3 subjects were experienced physical therapists in other areas who were also practicing in chest physical therapy on occasional weekend service. Group 4 subjects were new graduates. Ten patients were evaluated by each group of 4 physical therapists using a teaching stethoscope with one diaphragm/bell and four pairs of earpieces. The education session consisted of discussion of the adoption of standardized nomenclature and education on proper technique and interpretation of auscultation. Interrater reliability was assessed before and after the education session using kappa (kappa) values. Comparisons were made between kappa values before and after the education session to determine the effect of education and between groups to determine the effect of clinical experience. RESULTS: The kappa values before the education session were low, indicating poor reliability in detecting specific abnormal sounds (kappa = -.02-.59). Group 1 (seniors in respiratory therapy) and group 4 (new graduates) demonstrated the greatest reliability levels. The lowest kappa values were observed for detecting and categorizing the quality of breath sounds (normal, absent, bronchial, or decreased) (kappa = -.02-.25). Following the education session, there was a general improvement in reliability (kappa = -.30-.77), especially for group 3 (specialists in other areas). The most improvement was noted for the detection of the quality of breath sounds (kappa = .08-.50). CONCLUSION AND DISCUSSION: Reliability of auscultation was poor to fair, in general, before the education session. There was a definite improvement in reliability after the education session. There was no clear effect of clinical experience on reliability, and the agreement among observers appeared to depend on the abnormal lung sound present. Limitations of this study and recommendations for future research are discussed. [Brooks D, Thomas J. Interrater reliability of auscultation of breath sounds among physical therapists.

Auscultation

Goniometric reliability in a clinical setting. Elbow and knee measurements.

Reliability of goniometric measurements has been examined only under standardized conditions and usually with healthy subjects. The purpose of this study was to assess goniometric reliability in a clinical setting. The reliability of goniometric measurements of passive elbow and knee positions was assessed using patients as subjects. The effect of using the means of repeated measurements and the interdevice reliability of three common goniometers were also examined. Results showed that intratester reliability for flexion and extension of the knee and the elbow joints was high (r = .91 to .99). Intertester reliability was also high (r = .88 to .97) for these measurements except for measurements of knee extension (r = .63 to .70). Although previous investigators have suggested that using the means of multiple measurements improves reliability, our data indicate that this procedure never improves the correlation coefficient more than .12. The reliability was similar for all three devices. The results of this study indicate that for the knee and elbow joints, goniometric measurements performed in a clinical setting can be highly reliable. The method described in this study provides a simple protocol that can be used clinically to investigate goniometric reliability.

Elbow Joint

Reliability of individual diagnostic criterion items for psychoactive substance dependence and the impact on diagnosis.

OBJECTIVE: Reliability of diagnostic criterion items for psychoactive substance dependence and the impact of each on the reliability of the diagnosis were analyzed. METHOD: As part of a reliability study for a new interview developed for the multisite Collaborative Study on the Genetics of Alcoholism (COGA), data were collected from both within-center and across centers. The impact of each diagnostic item on the reliability of the substance dependence diagnosis was studied by forcing each item to be reliable one at a time and recomputing the kappa statistic for the diagnosis. RESULTS: Findings indicated that the majority of individual diagnostic criterion items were reliable; 87% and 81% were in the fair or better range of reliability for the within- and cross-center studies, respectively. Individual kappa estimates were statistically similar for the two studies. Reliability findings for two classes of substance, alcohol and cocaine, were good, while those for stimulants were less satisfactory. CONCLUSIONS: Forcing items one at a time to be reliable did not affect reliability of the overall substance dependence diagnosis, because more than one criterion item changed from Time 1 to Time 2. Because no single item was influential, weighting criteria equally, as is done in the DSM and ICD classification systems, appears to be a reasonable approach.

Adolescent

Reliability analysis in therapeutic research: practice and procedures.

Twenty studies examining the reliability of assessment devices and outcome measures in therapeutic research were reviewed and analyzed. The 20 investigations contained 215 quantitative reliability values published in either the American Journal of Occupational Therapy or Physical Therapy during the past 5 years. The reliability studies were classified as interrater, intrarater, test-retest, or internal consistency. Examination of interrater reliability accounted for 41% of all reported reliability values. Studies published in Physical Therapy were more likely to be concerned with test-retest reliability, whereas studies published in the American Journal of Occupational Therapy more often focused on interrater reliability. Examination of the data revealed that the intraclass correlation coefficient (ICC) was the most frequently reported estimate of reliability, accounting for 57% of all reported reliability coefficients. Further review of the results indicated that Pearson product-moment correlations and percentage of agreement indexes accounted for 22% of all reliability values reported in the studies examined. The Pearson product-moment correlation measures association or covariation among variables, but not agreement, and percentage agreement indexes do not correct for chance agreement. The argument is made that product-moment correlations and percentage agreement indexes are inadequate measures of interrater, intrarater or test-retest agreement. They should be used and interpreted with caution.

Humans

An examination of reliability in developmental research.

The purpose of this investigation was to examine quantitative methods used to determine reliability in developmental research. Procedures used to compute reliability estimates in 30 studies published in three developmental journals were examined. Four types of reliability studies were identified and analyzed. These included interrater reliability, stability (test-retest and intrarater reliability), equivalence reliability, and internal consistency. Interrater reliability investigations were the most frequently reported in the developmental literature reviewed (45%). The Pearson product moment correlation (r) was the most commonly reported reliability statistic. The findings reveal that researchers in developmental pediatrics frequently analyze reliability data using the Pearson product moment correlation and interpret the results as indicating consensus (agreement) among raters or across instruments. The Pearson product moment correlation (r) provides information on covariation among variables but does not indicate agreement. Thus, the findings suggest that developmental researchers may be misinterpreting the statistical results of reliability investigations. The argument is made that the intraclass correlation coefficient (ICC) is a more appropriate method of analysis when the purpose of the research is to examine consensus.

Adolescent

Individual reliability of amplitude distribution in topographical mapping of EEG.

Whilst there is an accumulation of evidence suggesting that many quantitative EEG parameters show good stability and reliability, no previous study has considered whether the spatial distribution of EEG amplitude is reliable over time within a session. This study reports on the spatio-temporal reliability of EEG using data recorded from 24 subjects in a baseline condition with eyes open and also whilst performing a simple motor task. Both the internal stability and test-retest reliability for electrode parameters were comparable to previously published data. For most individuals, amplitude distribution was stable within each recording condition, but the test-retest reliability after 40 min was less good with the poorest reliability in the delta frequency band. Most subjects showed spatio-temporal reliability of less than 0.7 in at least one frequency band. In contrast, spatio-temporal reliability for the group average was good and exceeded 0.88 in all frequency bands. It is argued that the results indicate that reliability is insufficient to allow topographical comparisons for a single individual, but is more than adequate to allow group comparisons.

Adolescent

Reliability of radiographic assessment of acromial morphology.

The most widely used radiographic classification system for acromial morphology identifies three distinct acromial shapes: type I (flat), type II (curved), and type III (hooked). The purpose of this study was to measure the interobserver and intraobserver reliability of determinations of acromial morphology as defined by this system. Between 1990 and 1992, one hundred twenty-six supraspinatus outlet radiographs were obtained from 126 patients by technicians from Triangle Orthopaedic Associates in Durham, N.C. Six fellowship-trained shoulder surgeons independently reviewed each radiograph and classified it as type I, II, or III on the basis of established guidelines. Two surgeons classified each film a second time in random order. Analysis of variance was performed to obtain coefficients for interobserver and intraobserver reliability. Consensus ratings were then used to classify the 126 radiographs into consensus type I, consensus type II, or consensus type III groups. Percentages of type I, II, and III individual ratings within each consensus group were determined. The intraobserver reliability coefficient was 0.888, interpreted as good to excellent reliability. The interobserver reliability coefficient was 0.516, interpreted as poor to fair reliability. Of the 126 radiographs, 26 (20.6%) were rated as consensus type I, 76 (60.3%) were rated as consensus type II, and 24 (19.1%) were rated as consensus type III. The reliability of observer ratings was lowest when delineation between acromial types II and III was required. The low interobserver reliability makes comparisons of studies by different authors difficult to interpret and obscures the true incidence of acromial morphologic types. It also questions reported correlations between acromial type and shoulder pathologic conditions. It is concluded that a system that incorporates more objective classification criteria and acknowledges the continuous nature of acromial morphologic types may improve interobserver reliability and validate the system's use in making clinical and surgical judgments.

Acromion

The test-retest reliability of standardized instruments among homeless persons with substance use disorders.

OBJECTIVE: Standardized instruments are widely used to assess homeless persons, but basic data on their reliability and validity in these populations have not been available. The purpose of this study was to examine the reliability of standardized instruments used in a cooperative agreement on homeless persons with substance use disorder. METHOD: This study examined the 1-week test-retest reliability of the Alcohol Dependence Scale, the Addiction Severity Index and the Personal History Form, using 189 randomly selected subjects participating in a multisite study of services for homeless persons with alcohol and other drug abuse problems. In addition to scales and items, factors hypothesized to influence reliability related so subject, interviewer and setting were examined. RESULTS: Results showed substantial reliability for scale scores (> .60) but mixed reliability for individual items. Reliability was greater when items were factual and based on a recent time interval, and when subjects were interviewed in a protected setting. Higher reliability was also related to younger age, female gender, a first episode of homelessness and lower severity of psychiatric problems. CONCLUSIONS: Reliability should be examined in individual studies of homeless persons, and efforts should be made to minimize controllable sources of unreliability.

Adult

The reliability of goniometric measurements of active and passive wrist motions.

A reliability study was conducted to determine (a) the intrarater and interrater reliability of goniometric measurement of active and passive wrist motions under clinical conditions and (b) the effect of a therapist's specialization on the reliability of measurement. Randomly paired therapists performed repeated measurements of active and passive wrist motions in 48 subjects who had been referred to one of four occupational therapy or hand management clinics for evaluation and treatment. The data were analyzed with an intraclass correlation coefficient. A posteriori data analyses were performed to determine the effects of identified sources of error on the reliability of measurement. The results indicated that measurement of wrist motion by individual therapists is highly reliable and that intrarater reliability is higher than interrater reliability for all active and passive motions. Interrater reliability was generally higher among specialized therapists for reasons not immediately apparent from this study. With the exception of pain, identified sources of error were found to have surprisingly little effect on the reliability of measurement.

Adult

Reliability of clinical findings in temporomandibular disorders.

The aim of the present investigation was to study the interexaminer reliability of orthopedic tests and palpation techniques routinely used in the clinical diagnosis of disorders of the masticatory system. The tests were performed by a dentist and a physiotherapist, who both used the tests routinely when examining patients with temporomandibular disorders. Seventy-nine patients participated in this study. In the analysis, percentage agreement, intraclass correlation, and Cohen's kappa were used. The interexaminer reliability of the tests measuring maximal active mouth opening and registration of clicking during active mouth opening was high. The interexaminer reliability was fair for the tests measuring the intensity of pain during active movements and moderate for tests recording joint sounds (kappa = 0.47 to 0.59). There was high interobserver agreement on several items of the traction and translation tests, although the kappa values were low. The interexaminer reliability of the multitest scores for compression was substantial for joint sounds (kappa = 0.66) and fair for pain (kappa = 0.40). The interexaminer reliability of the multitest scores for muscle palpation and joint palpation was moderate (kappa = 0.51) and fair (kappa = 0.33), respectively. It can be concluded that most variables determined during active movements can be measured with satisfactory reliability, whereas variables for other tests are not measured with the same reliability on the basis of the kappa scores. The main symptoms of temporomandibular disorders can be evaluated reliably with multitest scores. It is recommended that clinicians calibrate their techniques regularly to improve the reliability of results in daily practice.

Adolescent

Reliability, Device Agreement and Validity of Load-Velocity Profiles: A Systematic Review with Meta-analysis.

BACKGROUND: For a valid one-repetition maximum (1RM) prediction via load-velocity (LV) relationships, high reliability and accuracy must be assumed. OBJECTIVE: Since individual study results indicate ambivalent prediction, this systematic review and meta-analysis was designed to provide a updated and comprehensive overview, extending knowledge about the validity and reliability of commercially available velocity sensors in Part I and the validity and reliability of velocity-based 1RM prediction models in Part II. METHODS: A systematic literature search was conducted in PubMed/MEDLINE, Web of Science, and Scopus. Validity and/or reliability studies or velocity-based 1RM prediction evaluations were included. Methodological quality was assessed using adapted COSMIN. The analysis was performed for intraclass correlation coefficient (ICC), Lin's concordance correlation coefficient (CCC), and Pearson's correlation coefficient (r). The review was preregistered in PROSPERO (CRD42025634595). RESULTS: Sixty-three studies were included for sensor validity and reliability and 38 for 1RM prediction models. Part I: Velocity sensors demonstrated good-to-excellent pooled validity and device agreement (ICC = 0.91-0.92 [0.83-0.97]; k = 55 and 439, respectively); intra- and inter-day reliability were classified as good to excellent with ICC = 0.90-0.91 [0.85-0.95] (k = 228 and 608, respectively), with sensor technology moderating the results. However, substantial heterogeneity and wide ranges of study-level estimates indicated considerable variability across moderators, linear position transducer (LPT) generally showing more consistent performance than inertial measurement units (IMU). Part II: Velocity-based 1RM prediction showed ICCs = 0.90 [0.83-0.94] (k = 124) and ICC = 0.91 [0.72-0.98] (k = 9); for reliability and validity, respectively. DISCUSSION: Commercial velocity sensors generally provide high relative validity and reliability. Results varied depending on exercise complexity, intensity, sensor technology, and modeling approach. While velocity-based 1RM prediction demonstrated high average validity, large heterogeneity in lower body exercises significantly biased the results. Furthermore, the dearth of measurement error and agreement analyses prohibits final conclusions. CONCLUSION: Therefore, velocity-based monitoring and 1RM prediction require cautious interpretation, as sensor- and exercise-specific evidence remains limited.

Load–velocity relationship

The reliability of distance run tests for children in grades K-4.

The purpose of this study was to determine test-retest reliability for the 1-mile, 3/4-mile, and 1/2-mile distance run/alk tests for children in Grades K-4. Fifty-one intact physical education classes were randomly assigned to one of the three distance run conditions. A total of 1,229 (621 boys, 608 girls) completed the test-retests in the fall (October), with 1,050 of these students (543 boys, 507 girls) repeating the tests in the spring (May). Results indicated that the 1-mile run/walk distance, as recommended for young children in most national test batteries, has acceptable intraclass reliability (.83 less than R less than .90) for both boys and girls in Grades 3 and 4, has minimal (fall) to acceptable (spring) reliability for Grade 2 students (.70 less than R less than .83), but is not reliable for children in Grades K and 1 (.34 less than R less than .56). The 1/2 mile was the only distance meeting minimal reliability standards for boys and girls in Grades K and 1 (.73 less than R less than .82). Results also indicated that reliability estimates remained fairly stable across gender and age groups from the fall to spring testing periods, with the exception of the noticeably improved values for Grade 2 students on the 1-mile run/walk test. Criterion-referenced reliability (P, percent agreement) was also estimated relative to Physical Best and Fitnessgram run/walk standards. Reliability coefficients for all age group standards were acceptable to high (.70 less than P less than .95), except for Fitnessgram standards for 5-year-old girls on the 1-mile test for both fall and spring and for 6-year-old boys and girls on the 1-mile test administered in the spring.

Child

Getting the story straight: evaluating the test-retest reliability of a university health history questionnaire.

This study was designed to establish the reliability of a health history questionnaire used as a screening tool for incoming university students. The authors used a test-retest design, with a test interval of 6 months, on a sample of medical and nursing students. The analysis focused on overall reliability of the questionnaire and reproducibility of specific items, based on question format. Questionnaire items of specific interest were those with dichotomous yes/no response options versus open-ended format questions, those using the words frequently or recently, or those that asked multiple questions. Demographic characteristics of the subjects were considered in the evaluation of reliability. Overall reliability of the questionnaire (93.6%) was above the anticipated level of 90%, and subject sex or program of study did not show any significant differences in reproducibility of responses. Although wording of questions did not affect item reliability, dichotomous format questions demonstrated a higher degree of reliability (96.4%) than the overall reliability of the questionnaire. Recommendations for enhancing the reliability of the questionnaire are based on item analysis and information gathered from interviews with subjects.

Adult

Goniometric reliability in a clinical setting. Subtalar and ankle joint measurements.

Measurements of the subtalar joint neutral (STJN) position and passive range of motion (PROM) of the ankle joint and the subtalar joint (STJ) are often part of a physical therapy evaluation. These measurements may be used in treatment planning, such as in the prescription of specialized shoes or orthoses. Therefore, reliability of these measurements, as they are obtained clinically, must be determined. The purpose of this study was to examine the reliability of measurements of the STJN position and of ankle and STJ PROM. To determine reliability, repeated measurements of the STJN position and of STJ PROM were taken on the involved feet of 43 patients with neurologic orthopedic disorders (including both feet of 7 patients), and measurements of ankle PROM (dorsiflexion and plantar flexion) were taken on 42 of these patients (including both feet of 7 patients). Intraclass correlation coefficients (ICCs) for intratester reliability ranged from .74 to .90 for ankle and STJ measurements. The ICCs for intertester reliability were .25 for measuring the STJN position, .32 for STJ inversion, and .17 for SJJ eversion. The ICCs for intertester reliability were .50 for ankle dorsiflexion and .72 for ankle plantar flexion. Goniometric measurements of the STJN position and of PROM of the ankle and STJ appear to be moderately reliable if taken by the same therapist over a short period of time. With the exception of ankle plantar flexion, these measurements cannot be considered to be reliable between therapists.

Adolescent

The reliability and functional validity of visual and semiautomatic sleep/wake scoring in the Møll-Wistar rat.

The present paper has three major objectives: first, to document the reliability of a published criteria set for sleep/wake scoring in the rat; second, to develop a computer algorithm implementation of the criteria set; and third, to document the reliability and functional validity of the computer algorithm for sleep/wake scoring. The reliability of the visual criteria was assessed by letting two raters separately score 8 hours of polygraph records from the light period from five rats (14,040 10-second scoring epochs). Scored stages were waking, slow-wave sleep-1, slow-wave sleep-2, transition type sleep and rapid eye movement (REM) sleep. The visual criteria had good interrater reliability [Cohen's kappa (kappa) = 0.68], with 92.6% agreement on the waking/nonrapid eye movement (NREM) sleep/REM sleep distinction (kappa = 0.89). This indicated that the criteria allow separate raters to independently classify sleep/wake stages with very good agreement. An independent group of 10 rats was used for development of an algorithm for semiautomatic computer scoring. A close implementation of the visual criteria was chosen. The algorithm was based on power spectral densities from two electroencephalogram (EEG) leads and on electromyogram (EMG) activity. Five 2-second fast Fourier transform (FFT) epochs from each EEG/EMG lead per 10-second sleep/wake scoring epoch were used to take the spatial and temporal context into account. The same group of five rats used in visual scoring was used to appraise reliability of computerized scoring. The computer score was compared with the visual score for each rater. There was a lower agreement (kappa = 0.57 and 0.62 for the two raters) than in interrater visual scoring [percent agreement 87.7 and 89.1% (kappa = 0.82 and 0.84) in the waking/NREM sleep/REM sleep distinction]. Subsequently, the computer scores of the raters were compared. The interrater reliability was better than the interrater reliability for visual scoring (kappa = 0.75), with 92.4% agreement for the waking/NREM sleep/REM sleep distinction (kappa = 0.89). The computer scoring algorithm was applied to data from a third independent group of rats (n = 6) from an acoustical stimulus arousal threshold experiment, to assess the functional validity of the scoring directly with respect to arousal threshold. The computer algorithm scoring performed as well as the original visual sleep/wake stage scoring. This indicated that the lower intrarater reliability did not have a significant negative influence on the functional validity of the sleep/wake score.

Algorithms