Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86Linked to original sources

Reliability and validity of the Functional Rating Index in older people with low back pain: preliminary report.

BACKGROUND AND AIMS: The Functional Rating Index (FRI) was developed to provide an assessment instrument which has not only clinical usefulness but also quantifies the patient's current state of pain and dysfunction in a reliable and valid manner for spinal conditions. There is no study on the FRI applied to older people with low back pain (LBP). The primary aim of this study was to evaluate the validity and reliability of the FRI in older people with LBP. METHODS: A total of 76 subjects aged 65 to 90 years with LBP, of which 37 were cognitively intact and were followed up on a second occasion, were assessed by the FRI, numeric rating scale (NRS), Roland Morris Questionnaire (RMQ) and spinal movement test. Reliability was assessed by statistical analysis of test results for test-retest and internal consistency. To assess construct validity, the FRI was compared with the RMQ. Concurrent validity was assessed using the NRS and spinal mobility test. RESULTS: The FRI demonstrated high internal consistency, with alpha=0.921 for test and alpha=0.901 for retest. Item-scale correlations were between 0.549-0.871. Test-retest correlation was 0.913 (p=0.000). There was very good construct validity between the FRI and the RMQ for test (r=0.663, p<0.000) and retest (r=0.603, p<0.000). The FRI showed high correlation with the NRS (r=0.701, p<0.000 for test; r=0.743, p<0.000 for retest) and no correlation with the spinal movement test (r=0.173, p=0.307 for test; r=0.024, p=0.888 for retest). CONCLUSIONS: In this preliminary report, the FRI appears to be easy to administer, seems to have significant validity and reliability, and may be useful in geriatric assessment of older people with LBP.

Aged↗

Measuring change in parotid gland size: test-retest reliability of a novel method.

BACKGROUND: There is currently no convenient method for measuring parotid gland hypertrophy, a common condition among patients with bulimia nervosa (BN) and anorexia nervosa (AN). OBJECTIVE: To develop a technique for reliably estimating change in parotid gland size. METHODS: A method for measuring facial width as a surrogate marker of parotid gland size was developed using calipers to measure between defined reference points located on the parotid gland region. The method was tested for reliability when performed by a single operator and used to determine face width measurements of 15 control subjects. RESULTS: Face width measurements were reliable when performed by a single operator. Face width measurements of control subjects ranged from 9.1 cm to 15.3 cm. DISCUSSION: The caliper method of measuring changes in parotid size is a novel method of measurement of parotid hypertrophy. It is quick, non-invasive and inexpensive and is highly reliable in the hands of a single operator.

Adolescent↗

Bonded retainers--clinical reliability.

Bonded retainers have become a very important retention appliance in orthodontic treatment. They are popular because they are considered reliable, independent of patient cooperation, highly efficient, easy to fabricate, and almost invisible. Of these traits, reliability is the subject of this clinical study. A total of 549 patients with retainers were analyzed with regard to wearing time, extension of the retainer, mean time between failures, operator, and age of patient. The average frequency of breakage or loss was 0.55 per retainer per year. This frequency was dependent primarily on the operator who bonded the retainer and on the extent of the retainer. If the upper canines were involved, reliability was lower. The majority of failures occurred during the first 3 to 6 months. The study showed that bonded retainers represent a highly efficient and reliable retention appliance suited to long-term use.

Adolescent↗

Inter-rater reliability of family history information on psychiatric disorders in relatives.

The family history method in psychiatric family studies is an important and necessary way of obtaining information on family members who are not available for personal interview. Studies on the validity of this method have shown that family history information on psychiatric disorders in relatives is neither accurate nor sensitive but highly specific. However, its inter-rater reliability has rarely been assessed, even though this is a prerequisite for adequate validity. In the present investigation we examined the inter-rater reliability of family history information obtained with a semi-structured and symptom-oriented interview. Forty informants were interviewed twice by two different raters within 3 and 20 days. The inter-rater reliability was found to be good for dementia (kappa=0.82, 95% CI=0.61-1.00), alcohol related disorders (kappa=0.93, 95% CI=0.80-1.00), for depressive disorders (kappa=0.72, 95% CI=0.42-1.00), anxiety disorders (kappa=0.75, 95% CI=0.41-1.00) and any psychiatric disorder (kappa=0.79, 95% CI=0.66-0.91). We concluded that the family history interview is a useful family study instrument that can be applied reliably by different raters for frequent psychiatric disorders.

Adult↗

Test-retest reliability of the computerized DSM-IV version of the Munich-Composite International Diagnostic Interview (M-CIDI).

The structure and content of the Munich-Composite International Diagnostic Interview (M-CIDI) for the assessment of DSM-IV symptoms, syndromes, and diagnoses is described along with findings from a test-retest reliability study. A sample of 60 community respondents were interviewed twice independently by trained interviewers with an average time interval of 38 days between investigations. Test-retest reliability was good for almost all specific DSM-IV core symptom questions and disorders examined, with kappa values ranging from fair for two diagnoses--bulimia (kappa 0.55) and generalized anxiety disorder (kappa 0.45)--to excellent (kappa above 0.72) for all other anxiety disorders and alcohol use disorders. Test-retest reliability for age of onset and time-related questions was fairly consistently high (intra-class correlation values of 0.79 or above), with one notable exception: the assessment of disorders with onset before puberty. We concluded that the M-CIDI is acceptable for respondents, efficient in terms of time needed for and ease of administration, and reliable in terms of consistency of findings over time periods of at least 1 month.

Adolescent↗

The rationale, development and reliability of a new screening psychiatric instrument.

BACKGROUND: This paper describes the rationale, development, reliability and validity of a new screening psychiatric instrument. METHOD: The instrument comprises 26 items that tap the cardinal features of main psychiatric categories as defined by ICD-10 and DSM-IV. These items were adapted from various structured and semi-structured diagnostic interviews that yield ICD-10 and DSM-IV psychiatric diagnoses. After a training course, 12 trainees and the trainer rated blindly the 26 items on 45 subjects (22 with psychopathology and 23 without). Inter-rater reliability coefficient (Kappa) was estimated between trainees and the trainer on each item of the instrument. The total score on the new instrument was then correlated with the total score on the Arabic Self Reporting Questionnaire (SRQ-20) and the Arabic version of the General Health Questionnaire (GHQ) in a random sample from the general population (n = 365). Logistic regression was utilised to estimate the power of the total score on the new instrument in discriminating between cases and non-cases as classified by the SRQ-20. RESULTS: Excellent levels of agreement (Kappa > 0.80) were found for all items except for obsession (Kappa = 0.65) and for depressed mood (Kappa = 0.70). Moderate correlations were found between the total score on the new instrument and total score on SRQ-20 (r = 0.69) and the total score on the Arabic GHQ (r = 0.7). The new instrument correctly classified 89% of subjects into cases and non-cases. CONCLUSIONS: The results of this study indicate that the new instrument is a highly reliable and valid screening instrument. The authors are now investigating its test-retest reliability and its procedural validity.

Adolescent↗

The reliability of case register diagnoses: a birth cohort analysis.

BACKGROUND: Few studies assess the reliability of case register diagnoses, despite their widespread use in psychiatric research. This study investigates case register diagnostic reliability in comparison to casenote derived diagnoses in a birth cohort. METHODS: Diagnostic information from the case register and casenotes of 449 individuals was extracted. The incident and lifetime register diagnoses were compared with those derived from the casenotes. RESULTS: Inter-rater reliability was good (kappa = 0.71). Agreement between casenote and incident register diagnosis was moderate (kappa = 0.52), as was agreement between casenote and register lifetime diagnosis (kappa = 0.58). Case register diagnoses were insufficiently accurate to stand alone. Case register diagnoses for organic disorder, schizophrenia, alcoholism, learning disability, personality disorder and transient or no psychiatric disorder were reliable enough for the case register to act as a useful screening instrument. The case register was not acceptable, even as a screening instrument, for the diagnoses of neurotic or affective disorders. CONCLUSIONS: Studies relying only on case register diagnoses may be flawed if diagnoses are not independently verified. National statistics derived from case register data, especially for neurosis and affective disorder, may be unreliable.

Cohort Studies↗

[Inter-observer reliability in the classification of thoraco-lumbar spinal injuries].

The purpose of a fracture classification is to help the surgeon to choose an appropriate method of treatment for each and every fracture occurring in a particular anatomical region. The classification tool should not only suggest a method of treatment, it should also provide the surgeon with a reasonably precise estimation of the outcome of that treatment. But to use a classification before its workability has been proved is inappropriate and can lead to confusion and more conflicting results. Any classification system should be proved to be a workable tool before it is used in a discriminatory or predictive manner. The radiographs of fourteen fractures of the lumbar spine were used to assess the interobserver reliability of the AO classification system. The radiographs and CT scans were reviewed in twenty-two hospitals experienced with spinal trauma. The mean interobserver agreement for all fourteen cases was found to be 67% (41-91%), when only the three main types (A, B, C) were used. The corresponding kappa value of the interobserver reliability showed a coefficient of 0.33 (range, 0.30 to 0.35). The reliability decreased by increasing the categories. For some injuries the interobserver reliability was found to be over 90% and also for the recommended therapeutic procedure there was an acceptable agreement. But the decision between an posterior approach alone or an additionally anterior procedure seems to be the most important question in treatment of spinal injuries at that time.

Accidents↗

Reliability of isokinetic ankle inversion- and eversion-strength measurement in neutral foot position, using the Biodex dynamometer.

This study was designed to investigate the intratester and intertester reliability of isokinetic ankle inversion and eversion-strength measurement in neutral foot position in healthy adults using the Biodex dynamometer. Twenty-five men and women performed five maximal concentric contractions at 60 and 180 degrees/s angular velocities. Two physicians tested each subject. The first physician applied the test four times, and the second physician three times. Reliability of peak torque was assessed by calculating the intraclass correlation coefficient (ICC). At both angular velocities, inversion strength was greater than eversion, and when the angular velocity was increased, inversion and eversion strength were decreased, as tested by both physicians. The first measurements of inversion and eversion strength of the first physician were significantly lower than the other measurements (p<0.01). The intratester ICCs for ankle inversion in healthy young adults were highly reliable (ICC 0.92-0.96), and for the eversion values ranged from 0.87 to 0.94. Intertester ICCs for ankle inversion and eversion peak torque values demonstrated a value of 0.95. Isokinetic tests of ankle inversion and eversion strength at 60 and 180 degrees/s angular velocities in neutral foot position for healthy adults are highly reliable with the Biodex dynamometer.

Adult↗

Two ankle joint laxity testers: reliability and validity.

Two test devices were manufactured to objectively measure ankle joint laxity: the dynamic anterior ankle tester (DAAT) and the quasi-static anterior ankle tester (QAAT). The primary aim was to analyse the reliability of both testers; The secondary aim was to assess validity in correlation with TELOS stress test and manual anterior drawer test. Twenty-four normal subjects and 14 patients 1 year after acute lateral ankle ligament injury were included. Both ankles were tested with the DAAT and QAAT by two different observers; one experienced orthopaedic surgeon performed the manual test; the TELOS stress X-rays were evaluated by one observer. Intra observer reliability for the DAAT varied between 0.81 and 0.94; for the QAAT between 0.71 and 0.94. Inter observer reliability for the DAAT varied between 0.84 and 0.94; for the QAAT between 0.76 and 0.82. Concurrent validity showed fair correlation between DAAT and QAAT for the first couple observers (0.71); however, a poor correlation was observed for the second couple (0.42). No significant correlations were found between neither DAAT and the TELOS and the manual test, nor QAAT and the TELOS and the manual test. In conclusion, reliability of both testers is high. Validity of the testers needs further investigation.

Adult↗

The validity and reliability of the Turkish version of the Fibromyalgia Impact Questionnaire.

This study was undertaken to translate and adapt the Fibromyalgia Impact Questionnaire (FIQ) into the Turkish language and investigate its validity and reliability for Turkish female fibromyalgia (FM) patients. After translation into Turkish, we administered the FIQ and Health Assessment Questionnaire (HAQ) to 51 women with fibromyalgia. As well as sociodemographic characteristics, the severity of relevant clinical symptoms, e.g., pain intensity, fatigue, and sleep disturbance, were assessed by visual analog scales. A tender point score (TPS) was calculated from tender points conducted by thumb palpation. Test-retest reliability, internal consistency, and concurrent and construct validities of FIQ were evaluated. Test-retest reliability and internal consistency were good at 0.81 and 0.72, respectively. Correlation between FIQ and HAQ scores was 0.43, which was low but statistically significant. Significant moderate correlations were obtained between the FIQ items and severity of clinical symptoms (0.63-0.77), except TPS, 0.31. The FIQ is a reliable and valid instrument for measuring functional disability in Turkish female FM patients.

Adult↗

Magnetic resonance imaging of shoulders with idiopathic adhesive capsulitis: reliability of measures.

The magnetic resonance imaging (MRI) findings in idiopathic adhesive capsulitis (AC) were compared with those of contralateral healthy shoulders and the reliability of measures assessed. Twenty-six consecutive patients (26 AC and 14 healthy shoulders) were prospectively assessed. The main measurements were thickness of the joint capsule and synovial membrane in the axillary recess and rotator interval in T1-weighted spin-echo sequence enhanced with intravenous (IV) gadolinium chelate (Gd-chelate). Reliability was studied by use of the intraclass correlation coefficient (ICC). The mean thickness of the axillary recess on the coronal plane was 9.0+/-2.2 mm in AC shoulders and 0.4+/-0.7 mm in healthy shoulders. The mean thickness of the rotator interval on the sagittal plane was 8.4+/-2.8 in AC shoulders and 0.6+/-0.8 mm in healthy shoulders. Interobserver reliability was good for the axillary recess, with ICC values of 0.84 for the coronal plane, and good for the rotator interval, with ICC values of 0.80 for the sagittal plane. MRI with IV Gd-chelate injection can show, with acceptable reliability, signal and thickness abnormalities of the shoulder joint capsule and synovial membrane in AC.

Adult↗

White matter hyperintensities and rating scales-observer reliability varies with lesion load.

BACKGROUND: Cerebral white matter hyperintensities (WMHs) are common in older people. Their presence correlates with cognitive decline and vascular risk factors. Various scales have been developed to quantify the amount and type of WMH, but with few observer reliability studies. We evaluated several scales in different cohorts to determine their observer reliability. METHODS: Two observers independently rated T2-weighted MR images from five groups (total n = 494: normal older subjects [97]; patients with minor stroke [221]; young insulin dependent diabetics [141]; maturity onset diabetics [10]; and hepatic encephalopathy [25]), using seven rating scales (Breteler, Fazekas, Longstreth, Mirsen, Shimada, Van Swieten and Wahlund). Inter-observer reliability was determined using Kappa statistics. RESULTS: Patients with maturity onset diabetes had the most WMHs and young insulin-dependent diabetics the least. Inter-observer reliability varied with the amount of WMH. In maturity onset diabetics (most WMHs) the weighted Kappas were: Breteler 0.74; Fazekas 0.89 and 0.72; Van Swieten 0.76 and 0.88; and in young insulin-dependent diabetics (least WMH): Breteler 0.3; Fazekas 0.2 and 0.24; Van Swieten 0.39 and 0.30. These findings were consistent across the groups. CONCLUSION: WMH rating scale performance varied with WMH prevalence, and hence with subject cohort. In patients with most WMHs the apparent better kappas may reflect a "ceiling effect" rather than true better agreement. These factors should be considered in studies where risk factors for, or associations with, the early development of WMHs are being determined.

Adult↗

The 5 min running field test: test and retest reliability on trained men and women.

The repeatability of the maximal aerobic velocity ( v(amax)) estimated using the 5 min running field test (5(RFT)) has been examined in an heterogeneous population of 132 subjects distributed in five groups considering their sporting activities, their competition levels and their physical fitness levels: among them were national and local runners, rugby players, and multi-sport women and men. To establish the test and retest reliability, all the subjects took part in the 5(RFT) twice within 3 weeks. After the normality of distributions had been assessed using a Kolmogorov-Smirnov's test, a Student's paired t-test showed no difference between the two trials in all groups except that of the national runners. A heterogeneous group was then constituted from the other subjects, and this took part in the reliability study. Intraclass correlation coefficients calculated from a one-way ANOVA on the performances achieved by each group in both tests ranged from 0.94 to 0.98. The standard errors of measurement (SEM) of the 5(RFT) ranged from 0.95% (13 m) to 1.89% (20 m), which correspond to errors of 0.15 km.h(-1) and 0.34 km.h(-1) in the v(amax), respectively. These results indicate that the 5(RFT) is reliable when used in homogeneous groups with various characteristics as well as in a heterogeneous population. Moreover, the results of this study have shown that the 5(RFT) is reliable for estimating v(amax) from only one trial, since the intraclass correlation coefficients for one trial ranged from 0.88 to 0.96, which is of particular interest to coaches. Nevertheless, further studies would be necessary to evaluate the repeatability of this test in other populations such as school children and adults of both sexes having different characteristics.

Adolescent↗

Comparison and reliability of two non-invasive acetylene uptake techniques for the measurement of cardiac output.

Comparison and reliability of two non-invasive acetylene uptake techniques for the measurement of cardiac output. Thirteen trained male cyclists performed CO2 rebreathing (CO2RB) at intensities from rest to 200 W, and open-circuit acetylene uptake (OpCirc) and single-breath acetylene uptake (SB) at intensities from rest to 300 W, with all procedures using 50 W increments. Oxygen consumption VO2 cardiac output Q and heart rate (HR), were measured at each stage, and the values for each variable were compared within each intensity to determine reliability of the measuring device. Both the OpCirc and SBs were shown to be reliable measures of cardiac output (r = 0.95 and 0.92, respectively) with decreasing coefficients of variation (CV) as intensity increased, and were similar to published data. The Q-VO2 relationship using the SB diverged from the regression line for OpCirc and CO2RB. Linear regression of the Q--VO2 relationship for CO2RB was y = 6.18 x VO2 + 2.59 for OpCirc was y = 6.12 x VO2 + 2.98 and for SB was y = 5.05 x VO2 + 3.76. The OpCirc and SBs were both shown to be reliable techniques for measuring cardiac output, comparable to previously reported cardiac output measurements, and suitable for use in exercise testing. However, the SB, requiring a constant, slow exhalation rate, made the procedure difficult to perform at higher exercise intensities.

Acetylene↗

Reliability of a cycling time trial in a glycogen-depleted state.

The aim of this study was to investigate the reliability of a protocol designed to simulate endurance performance in events of long duration (approximately 5 h) where endogenous carbohydrate stores are low. Seven male subjects were recruited (age 27 +/- 7 years, VO(2max) 66 +/- 5 ml/kg/min, W (max) 367 +/- 42 W). The subjects underwent three trials to determine the reliability of the protocol. For each trial subjects entered the laboratory in the evening to undergo a glycogen-depleting exercise trial lasting approximately 2.5 h. The subjects returned the following morning in a fasted state to undertake a 1-h steady-state ride at 50% W (max) followed by a time trial of approximately 40-min duration. Each trial was separated by 7-14 days. The trials were analysed for reliability of time to completion of the time trial using a coefficient of variation (CV), with 95% confidence intervals (data are mean +/- SD). The times to complete the three trials were 2,546 +/- 529, 2,585 +/- 490 and 2,568 +/- 555 s for trials 1, 2 and 3, respectively. The CV between trials 1 and 2 was 4.5% (95% CI 2.9-10.4%) and between trials 2 and 3, 3.8% (95% CI 2.4-9.9%). There was no difference in oxygen uptake, respiratory exchange ratio, carbohydrate oxidation, fat oxidation, plasma glucose concentration and plasma lactate concentration between the three trials. Therefore we can conclude that prior glycogen depletion does produce a reliable measure of performance with metabolic characteristics similar to ultraendurance exercise.

Adult↗

Oropharyngeal scintigraphy: a reliable technique for the quantitative evaluation of oral-pharyngeal swallowing.

A valid and reliable technique to quantify the efficiency of the oral-pharyngeal phase of swallowing is needed to measure objectively the severity of dysphagia and longitudinal changes in swallowing in response to intervention. The objective of this study was to develop and validate a scintigraphic technique to quantify the efficiency of bolus clearance during the oral-pharyngeal swallow and assess its diagnostic accuracy. To accomplish this, postswallow oral and pharyngeal counts of residual for technetium-labeled 5- and 10-ml water boluses and regional transit times were measured in 3 separate healthy control groups and in a group of patients with proven oral-pharyngeal dysphagia. Repeat measures were obtained in one group of aged (> 55yr) controls to establish test-retest reliability. Scintigraphic transit measures were validated by comparison with radiographic temporal measures. Scintigraphic measures in those with proven dysphagia were compared with radiographic classification of oral vs. pharyngeal dysfunction to establish their diagnostic accuracy. We found that oral ( p = 0.04), but not pharyngeal, isotope clearance is swallowed bolus-dependently. Scintigraphic transit times do not differ from times derived radiographically. All scintigraphic measures have extremely good test-retest reliability. The mean difference between test and retest for oral residual was -1% (95% CI -3%-1%) and for pharyngeal residual it was -2% (95% CI -5%-1%). Scintigraphic transit times have very poor diagnostic accuracy for regional dysfunction. Abnormal oral and pharyngeal residuals have positive predictive values of 100% and 92%, respectively, for regional dysfunction. We conclude that oral-pharyngeal scintigraphic clearance is highly reliable, bolus volume-dependent, and has a high predictive value for regional dysfunction. It may prove useful in assessment of dysphagia severity and longitudinal change.

Adolescent↗

Interrater and intrarrater reliability of the exeter dysphagia assessment technique applied to healthy elderly adults.

The purpose of this study was to evaluate the inter- and intrarater reliabilities of the Exeter Dysphagia Assessment Technique in a sample of elderly adults. This procedure uses noninvasive methods to record aspects of oral motor efficiency and synchronization of respiration during swallowing with the aid of specially developed equipment. Changes in the direction of nasal air flow, time of lip or tongue/spoon contact, and the time/frequency of swallow sounds are monitored and analyzed. Seventy records were evaluated independently by three trained assessors on three consecutive occasions. Interrater reliability was found to be good to very good for five of the respiratory variables assessed and moderate for the sixth. Interrater agreement was also very good for three of the timed oropharyngeal events assessed and moderate for the fourth. Intrarater reliability was very good for the same five respiratory variables and moderate for the sixth. Intrarater agreement was also very good for three of the timed oropharyngeal events and moderate for the fourth. Repeat evaluations of these records showed that agreement between and within raters concerning the sixth respiratory variable was improved substantially when the charts were examined in an enlarged form that provided improved resolution. We conclude that the majority of variables monitored by the Exeter Dysphagia Assessment Technique can be evaluated very reliably.

Aged↗