Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Test-retest reliability of a functional MRI working memory paradigm in normal and schizophrenic subjects.

OBJECTIVE: Repeated functional magnetic resonance imaging (fMRI) studies of schizophrenic subjects may identify brain activity changes in response to interventions. To interpret the findings, however, it is crucial to know the test-retest reliability of the measures used. METHOD: The authors scanned seven normal subjects and seven schizophrenic subjects on two occasions during performance of a working memory task. They quantified the reliability of task performance and brain activation. RESULTS: In both groups, task performance was reliable, and all a priori regions were activated in group-averaged test and retest data. In individual schizophrenic subjects, however, indices of cognitive activation were not reliable across sessions. Normal subjects showed reasonable reliability of activation. CONCLUSIONS: Even given reliable task performance, stable clinical status, and a stable pattern of group-averaged activation, individual subjects showed unreliable brain activation. This suggests that repeated fMRI studies of schizophrenia should control for sources of variation, both artifactual and intrinsic.

Adult↗

Reliability of the services assessment for children and adolescents.

OBJECTIVE: This study examined the test-retest reliability of a new instrument, the Services Assessment for Children and Adolescents (SACA), for children's use of mental health services. METHODS: A cross-sectional survey was undertaken at two sites. The St. Louis site used a volunteer sample recruited from mental health clinics and local schools. The Ventura County, California, site used a double-blind, community-based sample seeded with cases of service-using children. Participating families completed the SACA and were retested within four to 14 days. The reliability of service use items was calculated with use of the kappa statistic. RESULTS: The SACA- Parent Version had excellent test-retest reliability for both lifetime service use and previous 12-month use. The SACA also had good to excellent reliability when administered to children aged 11 and older for lifetime and 12-month use. Reliability figures for children aged nine and ten years were considerably lower for lifetime and 12-month use. The younger children's responses suggested that they were confused about some questions. CONCLUSIONS: This study demonstrates that parents and older children can reliably report use of mental health services by using the SACA. The SACA can be used to collect currently unavailable information about use of mental health services.

Adolescent↗

Reliability of global assessment of functioning ratings made by clinical psychiatric staff.

OBJECTIVE: In the Swedish psychiatric care system, systematic follow-up of clinical work with patients is becoming a part of regular service, and a number of care providers are using the Global Assessment of Functioning (GAF) to measure outcomes. This study investigated the reliability of the GAF and analyzed certain factors that affect measurement errors when the scale is used by regular psychiatric staff. METHODS: Eighty-one raters from various psychiatric outpatient clinics rated eight case vignettes. Interrater reliability was assessed by using intraclass correlation coefficients (ICCs), and factors associated with reliability were analyzed by using raters' unique residual values. RESULTS: The results showed that staff who are responsible for assessing first-time patients at outpatient psychiatric clinics and making diagnoses are using the GAF with satisfactory reliability (ICC(1,1)=.81). The factors associated with reliability were raters' subjective attitude toward the GAF and motivation to use the scale and other measurement instruments in psychiatry. CONCLUSIONS: GAF ratings made by an individual rater can be used to measure changes and outcomes at the group level. However, the measurement error is too large for assessment of change for an individual patient, in which case it might be necessary to use several raters. If raters are positively inclined to use rating instruments, measurement errors are minimized and reliability is maximized.

Diagnostic and Statistical Manual of Mental Disord↗

Reliability and validity of the substance abuse outcomes module.

OBJECTIVE: The study sought to determine the validity and reliability of the Substance Abuse Outcomes Module (SAOM), a self-report tool designed to assess patient characteristics, process of care, and outcomes of care, using a minimum amount of information, in order to improve treatment. METHODS: A longitudinal field test (baseline and three-month follow-up) compared the SAOM to seven other research instruments in the assessment of 100 substance-abusing patients who were entering a new treatment episode. Quota samples of patients were drawn from two private inpatient substance abuse treatment facilities and an outpatient methadone clinic. The study's primary outcome measures were diagnostic accuracy, internal and test-retest reliability of key constructs, concurrent and predictive validity, and sensitivity to change. Cronbach's alpha coefficients were calculated to examine internal consistency and reliability. Intraclass correlation coefficients and kappa coefficients were used to examine test-retest reliability. Concurrent validity of outcomes measures was examined with Pearson or Spearman correlation coefficients and chi square and kappa statistics. Changes between baseline and follow-up were examined as a function of case-mix measures with ordinary least-squares multiple regression. Sensitivity to change was examined by calculating effect size scores. RESULTS: The SAOM had high internal consistency and a high level of agreement with research diagnoses at baseline and follow-up. The SAOM was found to be highly reliable, to have very strong validity, and to be sensitive to clinical change. CONCLUSIONS: The SAOM appears to be a reasonably reliable and valid self-report instrument when used to monitor substance abuse treatment among patients with a primary substance use diagnosis.

Adolescent↗

Test-retest reliability study of a new improved Leg-O-meter, the Leg-O-meter II, in patients suffering from venous insufficiency of the lower limbs.

The objective of this study was to evaluate the interobserver and test-retest reliability of the new improved Leg-O-Meter, the Leg-O-Meter II, an instrument designed to measure leg circumference. The new Leg-O-Meter consists of a tape measure fixed to a stand attached to a small board on which the patient is in standing position. Only the left limb is measured. For this study the tape measure of the Leg-O-Meter was fixed at 13 cm from the board. Subjects were recruited from patients consulting the phlebology clinic of Hopital St-Michel, Paris, France. Thirty-nine patients were asked to participate in the test phase and a subsample of 20 patients were asked to participate in addition to a retest phase 10 minutes after their first measurement. Patients were asked to enter a closed room where four independent and blinded observers consecutively took measurements of their left calf with the Leg-O-Meter II. Twenty patients were also asked to come back 10 minutes later for a second round of measurements. While waiting, patients were seated. Variables collected included leg circumference, presence of edema, clinical presentation, and venous insufficiency treatment history. The order of the observers was randomized between patients. Under the assumption of a two-way random effects model, an intraclass correlation coefficient (ICC) was used to determine the reliability of a measure with the Leg-O-Meter II as well as the test-retest reliability. The interobserver and test-retest reliabilities of the Leg-O-Meter II were 98.28% [96.90%, 100.00%] CI95% and 95.90% [92.00%, 100.00%] CI95%, respectively. The Leg-O-Meter II has higher interobserver reliability and is easier to manipulate than the previous version. In addition, it has substantive test-retest reliability.

Anthropometry↗

The reliability of muscle function analysis using different methods of stimulation.

BACKGROUND: The purpose of our study was to determine the reliability of nonvolitional muscle function analysis (MFA) by determining the day-to-day and within-day reliability of conventional electrical stimulation and a newer, magneto-electrical stimulation method, using standard laboratory methodology. METHODS: Ten healthy, human immunodeficiency virus-negative adult men volunteered as subjects. MFA consisted of measuring the maximal relaxation rate, for magneto-electrical stimulation at 1 Hz and conventional electrical stimulation at 20 Hz, and force-frequency ratios using conventional electrical stimulation at 10 Hz:20 Hz and 10 Hz:50 Hz. Within-day and day-to-day reliability were determined by calculating the coefficient of variation (CV) for all subjects. RESULTS: Maximal relaxation rate using magneto-electrical stimulation had a significantly lower CV compared with the other nonvolitional MFA methods (p = .002). CONCLUSIONS: Maximal relaxation rate using magneto-electrical stimulation was more reliable and technically easier than the other muscle function parameters examined. However, the day-to-day CV of muscle function parameters is larger than traditional nutrition assessment techniques. Development within the field should strive to improve testing techniques so that the reliability of MFA will allow definition of a range of normal values against which an individual's value can be compared. Until this is available, the precision and reliability of MFA restrict its use to research and population studies.

Adult↗

Functional assessment of nursing home patients. Reliability and relevance.

Functional assessment of patients for resource allocation or staffing requires a higher level of interrater reliability than functional assessment in normal clinical settings. Although many functional assessment instruments are available, interrater reliability of these items has frequently not been reported. An assessment instrument based on the Long Term Care Minimum Data Set format was used for 290 patients in six wards in two Veterans Administration nursing homes. Each patient was assessed independently by two nurse caregivers to obtain reliability information. Different reliability measures yielded differing evaluations of the reliability of the instrument. Absolute agreement rates combined with Kendall's tau-b were most useful in deciding on the reliability of items in the instrument.

Activities of Daily Living↗

Are reliability, reproducibility and validity the correct terms to assess the correctness of dietary studies?

Nutritional studies often use the terms reliability, reproducibility and validity to indicate the correctness of the study. These terms do not appear to have a universal meaning to all researchers. The components of a dietary study are the input, the data collection instrument and the compiled data. Frequently the data collection questionnaire/tool/instrument is tested for reliability, reproducibility or validity. The data collection questionnaire/tool/instrument is simply a structure, a vehicle for gathering data. An argument is presented that demonstrates the reasons that such a structure cannot be tested for reliability, reproducibility or validity. The logical approach to the use of the terms reliability, reproducibility and validity is presented. Reliability refers to the input component of the study, reproducibility may or may not lead to strengthening the study and validity refers to the truthfulness of the database generated. Validity must be derived from reliable and reproducible data.

Diet Surveys↗

The reliability of assessing the appropriateness of requested diagnostic tests.

Despite a poor reliability, peer assessment is the traditional method to assess the appropriateness of health care activities. This article describes the reliability of the human assessment of the appropriateness of diagnostic tests requests. The authors used a random selection of 1217 tests from 253 request forms submitted by general practitioners in the Maastricht region of The Netherlands. Three reviewers independently assessed the appropriateness of each requested test. Interrater kappa values ranged from 0.33 to 0.42, and kappa values of intrarater agreement ranged from 0.48 to 0.68. The joint reliability coefficient of the 3 reviewers was 0.66. This reliability is sufficient to review test ordering over a series of cases but is not sufficient to make case-by-case assessments. Sixteen reviewers are needed to obtain a joint reliability of 0.95. The authors conclude that there is substantial variation in assessment concerning what is an appropriately requested diagnostic test and that this feedback method is not reliable enough to make a case-by-case assessment. Computer support maybe beneficial to support and make the process of peer review more uniform.

Decision Support Systems, Clinical↗

Variability and reliability of joint measurements.

The purpose of this study was to determine the variability and reliability of joint measurements as carried out by three physician observers. The intratester variation and reliability of nine different joint measurements was determined in eight healthy subjects. The measurements were taken in eight sessions by each tester. In this population also the intertester variation and reliability was determined by the three observers. This was also done in a population of middle-aged athletes over a period of 2.5 years. The results indicate that it is difficult to show either an improvement or worsening of a joint motion of less than 5 degrees to 10 degrees for most joints measured by the same tester. The intertester variation is not consistent over a longer period of time, so differences between observers during long-term studies cannot be corrected on the basis of a single study at a single point in time. The reliability of all nine joint measurements is not very high, but is probably sufficient if the results are used to compare groups within a single population and for large studies with experienced observers. Because the reliability strongly depends on the interindividual variation, it is preferable to determine the reliability for each study population.

Adult↗

Diagnostic interviewing with children: the use and reliability of the diagnostic coding form.

There have been few attempts to standardize assessment methods in Child Psychiatry. This paper describes a semi-structured approach to diagnostic interviewing of the child. Thirty-four children six to 13 years of age, and their parents, were interviewed two weeks apart by two different psychiatrists. A diagnostic coding form consisting of 29 clinical symptom items, eight summary items, and nine positive health ratings was used. Three diagnostic items were also included: "severity of clinical condition," "probability of disorder," and "adjustment status." Twelve of the Time 2 interviews with the child and parent were videotaped and rated by three different psychiatrists. Results indicated that summary items had higher reliability than individual symptom items and the three diagnostic items had the highest reliability, suggesting reliability is better for broad classes of behaviour. Interrater reliability was higher for the face-to-face rating than videotaped ratings. This suggests first that face-to-face interviews are reasonably stable over a two week period and second, since videotaped ratings had lowest reliability on items that depended on inferences about the child's feedlings, an important source of variance in assessment may be the clinician's ability to empathize with the child and draw inferences about internal feeling-states. It was concluded that this interview schedule can be a part of routine clinical practice. It ensures a reasonably standard, yet flexible and reliable approach to diagnostic interviewing.

Adaptation, Psychological↗

Ankle alignment on lateral radiographs. Part 2: reliability and validity of measures.

BACKGROUND: In ankles with end-stage osteoarthritis or after total ankle replacement (TAR), radiographic landmarks based on joint surface morphology usually are obscured and inadequate for measurement. Two methods for quantifying anteroposterior tibial-talar alignment without relying on those landmarks were identified in a corollary cadaver-based study. This study aimed to verify reliability and validity of those candidate measures. METHODS: On clinical radiographs of 33 nonarthritic and 35 arthritic ankles, the anteroposterior tibial-talar alignment was quantified by the two methods; the tibial-axis-to-talus ratio (T-T ratio: the ratio into which the midlongitudinal axis of the tibial shaft divides the longitudinal talar length) and the posterior-tibial-line-to-talus ratio (P-T ratio: a similar ratio, but using the posterior longitudinal line along the tibial shaft). Two observers performed every measurement twice to evaluate intraobserver and interobserver reliability of the candidate measures. For nonarthritic ankles, the anteroposterior tibial-talar alignment was further determined by a control measure that directly quantified orientation of the talar dome relative to the tibial shaft. Correlation of the T-T and P-T ratios with the control measure was then evaluated for validity. RESULTS: Measurement of the T-T ratio with arthritic ankles was highly reproducible with the coefficients of determination (R(2)) greater than 0.95, for either interobserver or intraobserver. Correlation between this measure and the control measure was supported (R(2) = 0.60, p < 0.0001). Reliability of the P-T ratio also was strong (R(2) > 0.91), although both reliability and validity of this measure were relatively inferior to the T-T ratio. CONCLUSIONS: The T-T ratio reliably and validly described the anteroposterior tibial-talar alignment on clinical radiographs, regardless of the condition of ankle joint surface. This measure appears to be a reliable radiographic measure for determining the magnitude of anteroposterior talar subluxation in ankles with articular degeneration or after TAR and can facilitate clinical investigations.

Adolescent↗

Interobserver and intraobserver reliability of two classification systems for intra-articular calcaneal fractures.

BACKGROUND: For a fracture classification to be useful it must provide prognostic significance, interobserver reliability, and intraobserver reproducibility. Most studies have found reliability and reproducibility to be poor for fracture classification schemes. The purpose of this study was to evaluate the interobserver and intraobserver reliability of the Sanders and Crosby-Fitzgibbons classification systems, two commonly used methods for classifying intra-articular calcaneal fractures. METHODS: Twenty-five CT scans of intra-articular calcaneal fractures occurring at one trauma center were reviewed. The CT images were presented to eight observers (two orthopaedic surgery chief residents, two foot and ankle fellows, two fellowship-trained orthopaedic trauma surgeons, and two fellowship-trained foot and ankle surgeons) on two separate occasions 8 weeks apart. On each viewing, observers were asked to classify the fractures according to both the Sanders and Crosby-Fitzgibbons systems. Interobserver reliability and intraobserver reproducibility were assessed with computer-generated kappa statistics (SAS software; SAS Institute Inc., Cary, North Carolina). RESULTS: Total unanimity (eight of eight observers assigned the same fracture classification) was achieved only 24% (six of 25) of the time with the Sanders system and 36% (nine of 25) of the time with the Crosby-Fitzgibbons scheme. Interobserver reliability for the Sanders classification method reached a moderate (kappa = 0.48, 0.50) level of agreement, when the subclasses were included. The agreement level increased but remained in the moderate (kappa = 0.55, 0.55) range when the subclasses were excluded. Interobserver agreement reached a substantial (kappa = 0.63, 0.63) level with the Crosby-Fitzgibbons system. Intraobserver reproducibility was better for both schemes. The Sanders system with subclasses included reached moderate (kappa = 0.57) agreement, while ignoring the subclasses brought agreement into the substantial (kappa = 0.77) range. The overall intraobserver agreement was substantial (kappa = 0.74) for the Crosby-Fitzgibbons system. CONCLUSIONS: Although intraobserver kappa values reached substantial levels and the Crosby-Fitzgibbons system generally showed greater agreement, we were unable to demonstrate excellent interobserver or intraobserver reliability with either classification scheme. While a system with perfect agreement would be impossible, our results indicate that these classifications lack the reproducibility to be considered ideal.

Ankle Injuries↗

The Foot Function Index for measuring rheumatoid arthritis pain: evaluating side-to-side reliability.

The Foot Function Index is a validated and reliable instrument for measuring foot pain, disability, and activity restriction in patients with rheumatoid arthritis. For the purposes of orthopaedic studies in which one foot serves as an internal control, we assessed the side-to-side reliability of the seven-question Foot Function Index pain subscale. Thirty patients with rheumatoid arthritis completed visual analog scale pain questionnaires for both feet on two occasions 8 days apart. Internal reliability of the scale was high, with Cronbach's alphas ranging from 0.94 to 0.96, suggesting good left versus right discriminatory abilities. Principal component factor analysis segregated the questions into two large clusters containing predominantly either left or right foot items. Intraclass correlation coefficients were examined for test-retest reliability (separated by side) and for side-to-side reliability (separated by the day of test). The resultant intraclass correlation coefficients were nearly equivalent, ranging from 0.79 to 0.89. Generalizability analysis yielded similar results. Intraclass correlation coefficients and generalizability analysis demonstrate that the majority of variation is best explained by the differences within subjects or between subjects rather than by test-retest or side-to-side differences. We recommend the Foot Function Index as a reliable measurement scale for use in orthopaedic interventional trials.

Adult↗

Test-retest reliability of standard and emotional stroop tasks: an investigation of color-word and picture-word versions.

Previous studies have examined the reliability of scores derived from various Stroop tasks. However, few studies have compared reliability of more recently developed Stroop variants such as emotional Stroop tasks to standard versions of the Stroop. The current study developed four different single-stimulus Stroop tasks and compared test-retest reliabilities. The four Stroop tasks included two standard Stroop tasks (color-word and picture-word) as well as two emotional Stroop tasks (color-word and picture-word). The four Stroop tasks were administered on two occasions, separated by 1 week, to 28 undergraduate students. Test-retest reliability coefficients were high for standard and emotional Stroop tasks when reliability was measured using response latencies alone. However, test-retest coefficients were unacceptably low when reliability estimates were calculated using difference scores. The findings have important implications for clinical and experimental use of standard and emotional Stroop tasks.

Adolescent↗

Research diagnostic criteria for temporomandibular disorders: a calibration and reliability study.

The aim of this study was to investigate the reliability between different examiners when using the axis I of the Research Diagnostic Criteria for Temporomandibular Disorders (RDC/TMD). The hypothesis was that the standardized RDC/TMD examination protocol enables calibrated examiners to evaluate all examination items reliably. After calibration training by the RDC/TMD calibration team including the calibration of palpation pressure and the performance of the standardized examination protocol, four examiners, blinded to the patients' medical histories examined 24 subjects in a randomized sequence. One experienced examiner was the standard (hierarchical calibration). The recorded measurements strictly followed the RDC/TMD. Intraclass correlation coefficients (ICC), bias and precision were calculated to estimate interrater reliability. Acceptable (0.75 > or = CC > 0.4) to excellent (ICC > 0.75) reliability was found for 20 of the 23 (87%) examinations. Only sub-retromandibular muscle palpation and joint sound vibration recordings on lateral excursion showed poor-results (ICC < or = 0.4). The RDC/TMD examination protocol enables calibrated examiners to perform most (87%) examination items with satisfactory reliability. Therefore multi-site studies based on the RDC/TMD examination protocol may become feasible, keeping in mind the unsatisfactory reliability of 13% of the items (clicking during laterotrusion to the ipsilateral side, palpation of the posterior and submandibular region).

Adolescent↗

The validity and reliability of the global index of safety (GIS).

OBJECTIVE: The global index of safety (GIS) is an adverse event (AE) based instrument designed to evaluate the safety profile of drugs. This paper presents the evaluation of the inter-rater reliability and validity of a 94-item GIS for antipsychotics through Rasch analysis. RESEARCH DESIGN AND METHODS: A total of 194 psychiatrists participating in an outpatient pharmacoepidemiologic study of olanzapine in schizophrenia rated the severity that each AE would have on a 5-point scale. Reliability was determined through a paired comparison design involving the new independent ratings of 101 different psychiatrists participating in another study of olanzapine in acute inpatient units. Spearman's, Pearson's and Intra-class correlation (ICC) coefficients were used to estimate the inter-rater reliability of the AE weights. Validity was analyzed through the Rasch rating scale model. RESULTS: Reliability coefficient estimates were excellent (Spearman = 0.99, Pearson = 0.99, ICC = 0.98), supporting the inter-rater reliability of the item weights. Through goodness-of-fit statistics and the investigation of the hierarchy of item calibrations, Rasch analysis confirmed the validity of the instrument. CONCLUSION: The data presented here on inter-rater reliability estimates of adverse events related to antipsychotic drugs indicate that GIS is a promising alternative for the evaluation of the safety profile of drugs.

Adverse Drug Reaction Reporting Systems↗

PSI-BLAST-ISS: an intermediate sequence search tool for estimation of the position-specific alignment reliability.

BACKGROUND: Protein sequence alignments have become indispensable for virtually any evolutionary, structural or functional study involving proteins. Modern sequence search and comparison methods combined with rapidly increasing sequence data often can reliably match even distantly related proteins that share little sequence similarity. However, even highly significant matches generally may have incorrectly aligned regions. Therefore when exact residue correspondence is used to transfer biological information from one aligned sequence to another, it is critical to know which alignment regions are reliable and which may contain alignment errors. RESULTS: PSI-BLAST-ISS is a standalone Unix-based tool designed to delineate reliable regions of sequence alignments as well as to suggest potential variants in unreliable regions. The region-specific reliability is assessed by producing multiple sequence alignments in different sequence contexts followed by the analysis of the consistency of alignment variants. The PSI-BLAST-ISS output enables the user to simultaneously analyze alignment reliability between query and multiple homologous sequences. In addition, PSI-BLAST-ISS can be used to detect distantly related homologous proteins. The software is freely available at: http://www.ibt.lt/bioinformatics/iss. CONCLUSION: PSI-BLAST-ISS is an effective reliability assessment tool that can be useful in applications such as comparative modelling or analysis of individual sequence regions. It favorably compares with the existing similar software both in the performance and functional features.

Information Storage and Retrieval↗