Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

Imprecision in orthodontic diagnosis: reliability of clinical measures of malocclusion.

The study examined the reliability among seven orthodontists in judging dental and facial aspects of malocclusion in a screening of elementary schoolchildren. Data included measures typically recorded during a clinical orthodontic examination: facial assessment of skeletal/incisor relationships and individual measures of morphologic malocclusion. Interexaminer reliability data were collected on 52 children. Pairwise comparisons between orthodontists were made using exact percent agreement and agreement within one category. Kappa statistics and one-sided Z-tests were used to evaluate observed agreement compared with agreement that would be expected by chance. Median Kappa statistics indicated that the reliability of maxillary and mandibular anteroposterior positions, incisor exposure, interlabial gap, and maxillary crowding was poor (K < 0.40). Acceptable reliability existed for mandibular anterior crowding, facial convexity, overbite, overjet, and molar classification (median Kappas ranged from 0.48 to 0.72). Excellent reliability existed only for evaluating the presence of a posterior crossbite (K = 0.79). The results caution that the language of clinical orthodontic diagnosis is imprecise.

Cephalometry↗

Validity and reliability of a new edge-based computerized method for identification of cephalometric landmarks.

OBJECTIVE: To evaluate the validity and inter- and intraexaminer reliability when on-screen landmarks are digitized manually or when these are computer-assisted by means of a new cephalometric software feature. MATERIALS AND METHODS: Twenty radiographs were digitized four times by two experienced orthodontists using a manual method and an edge-based algorithm that helps landmark identification by detecting the edges of anatomical structures. RESULTS: The computer-assisted method did not agree with manual digitization in 7 of 13 landmarks and 5 of 10 variables. With a tolerance of 0.5 mm or degrees, the two methods did not agree in cephalometric variables. Intraoperator reliability was improved for B point (x-axis), and Menton (x- and y-axis). It got worse for point A (y-axis). Interoperator reliability was improved for B point (x- and y-axis), Soft Labrale Inferior (x- and y-axis), Soft Pogonion (x-axis), and Menton (y-axis). It decreased for point A (y-axis). Intra- and interoperator reliability got better for only one cephalometric variable under study (SNB). CONCLUSIONS: The edge-locking feature seems to be a promising tool for increasing the reliability of on-screen cephalometric analysis. There seem to be difficulties in locating the appropriate edges when artifacts or soft tissue edges are located near the targeted landmark. The existence of very small, but systematic differences between the two digitization methods manifests the need for further improvement.

Algorithms↗

The effect of anchors and training on the reliability of perceptual voice evaluation.

Perceptual voice evaluation is a common clinical tool for rating the severity of vocal quality impairment. However, the evaluation process involves subjective judgment, and reliability is therefore a major issue that needs to be considered. When listeners are asked to judge the quality of a voice signal, they use their own internal standards as the references. These internal standards can be variable, as different individuals may have acquired different standards in prior situations. In order to improve the reliability of the perceptual voice evaluation process, external anchors and training are provided to counteract the effect of these internal standards. This study investigated to what extent the provision of anchors and a training program would improve the reliability of perceptual voice evaluation by naive listeners. The results show, in general, that anchors and training helped to improve the reliability of perceptual voice evaluation, especially in the rating of male voices. Furthermore, it was found that anchors made up of synthesized signals combined with training were more effective in improving reliability in judging perceptual roughness and breathiness than natural voice anchors.

Adult↗

Reliability issues and solutions for coding social communication performance in classroom settings.

PURPOSE: To explore the utility of time-interval analysis for documenting the reliability of coding social communication performance of children in classroom settings. Of particular interest was finding a method for determining whether independent observers could reliably judge both occurrence and duration of ongoing behavioral dimensions for describing social communication performance. METHOD: Four coders participated in this study. They observed and independently coded 6 social communication behavioral dimensions using handheld computers. The dimensions were mutually exclusive and accounted for all verbal and nonverbal productions during a specified time frame. The technology allowed for coding frequency and duration for each entered code. Data were collected from 20 different 2-min video segments of children in kindergarten through 3rd-grade classrooms. Data were analyzed for interobserver and intraobserver agreements using time-interval sorting and Cohen's kappa. Further, interval size and total observation length were manipulated to determine their influence on reliability. RESULTS: The data revealed interval sorting and kappa to be a suitable method for examining reliability of occurrence and duration of ongoing social communication behavioral dimensions. Nearly all comparisons yielded medium to large kappa values; interval size and length of observation minimally affected results. Implications The analysis procedure described in this research solves a challenge in reliability: comparing coding by independent observers of both occurrence and duration of behaviors. Results indicate the utility of a new coding taxonomy and technology for application in online observations of social communication in a classroom setting.

Adult↗

Reliability and validity characteristics of the Western Aphasia Battery (WAB).

The reliability and validity characteristics of the Western Aphasia Battery (WAB) are described. High internal consistency measures and high test-retest reliability argue for stability of the test both because its parts contribute to the composite index and because of its temporal reliability. Inter- and intrajudge reliability are both very high, suggesting consistent scoring within and between scorers. The WAB satisfies face- and content-validity criteria. Results from the WAB and the Neurosensory Center Comprehensive Examination for Aphasia (NCCEA) highly correlate, indicating good construct validity. WAB AQ scores and Raven's Coloured Progressive Matrices scores significantly correlate, suggesting that the language portions of the WAB are not totally independent from nonverbal functioning. WAB AQ scores reliably differentiate between aphasic and control groups, with only a small overlap for high functioning anomic aphasic subjects.

Adolescent↗

Individual differences and the reliability of 2F1-F2 distortion-product otoacoustic emissions: effects of time-of-day, stimulus variables, and gender.

Distortion product otoacoustic emissions (DPOAEs) measured from the ear canal can be a sensitive tool to detect changes in cochlear function over time. However, if multiple-measurement procedures are to be useful clinically, testing needs to be reliable and sources of variability within individuals should be known. Herein, the influence of time-of-day (TOD), stimulus frequency, stimulus sound pressure level (SPL), and gender were evaluated on 2f1-f2 DPOAE amplitude in 16 adult volunteers with normal hearing. The effects of oral temperature and resting-pulse rate were also assessed. This study demonstrated a TOD main effect, with a period approximating one cycle-per-day. The magnitude of this effect averaged less than one dB and was not dependent on stimulus (frequency or SPL) or participant variables (gender, oral temperature, or resting-pulse rate), nor was it synchronized to a particular point-in-time. Stimulus level and gender effects on DPOAEs across frequency were also observed. Using generalizability theory (GT), DP iso-level/frequency profiles (DPILFPs) were found to be reliable measures within-subjects over a contiguous 24-hour time period. Significant and reliable between-subject differences were also documented. This study demonstrates the influence of stimulus and participant variables, quantifies the within-subject reliability over a 24-hour time period, and confirms that significant and reliable between-subject differences exist on DPOAEs across frequency, SPL, and gender.

Acoustic Stimulation↗

The Alcohol, Smoking and Substance Involvement Screening Test (ASSIST): development, reliability and feasibility.

AIMS: The Alcohol, Smoking and Substance Involvement Screening Test (ASSIST) was developed for the World Health Organization (WHO) by an international group of substance abuse researchers to detect psychoactive substance use and related problems in primary care patients. This report describes the new instrument as well as a study of its reliability and feasibility. SETTING: The study was conducted at participating sites in Australia, Brazil, Ireland, India, Israel, the Palestinian Territories, Puerto Rico, the United Kingdom and Zimbabwe. Sixty per cent of the sample was recruited from alcohol and drug abuse treatment facilities; the remainder was drawn from general medical settings and psychiatric facilities. METHODS: The study was concerned primarily with test item reliability, using a simple test-retest procedure to determine whether subjects would respond consistently to the same items when presented in an interview format on two different occasions. Qualitative and quantitative data were also collected to evaluate the feasibility of the screening items and rating format. PARTICIPANTS: A total of 236 volunteer participants completed test and retest interviews at nine collaborating sites. Slightly over half of the sample (53.6%) was male. The mean age of the sample was 34 years and they had completed, on average, 10 years of education. RESULTS: The average test-retest reliability coefficients (kappas) ranged from a high of 0.90 (consistency of reporting 'ever' use of substance) to a low of 0.58 (regretted what was done under influence of substance). The average kappas for substance classes ranged from 0.61 for sedatives to 0.78 for opioids. In general, the reliabilities were in the range of good to excellent, with the following items demonstrating the highest kappas across all drug classes: use in the last 3 months, preoccupied with drug use, concern expressed by others, troubled by problems related to drug use, intravenous drug use. Qualitative data collected at the end of the retest interview suggested that the questions were not difficult to answer and were consistent with patients' expectations for a health interview. The data were used to guide the selection of a smaller set of items that can serve as the basis for more extensive validation research. CONCLUSION: The ASSIST items are reliable and feasible to use as part of an international screening test. Further evaluation of the screening test should be conducted.

Adult↗

Estimations of the efficacy and reliability of paternity assignments from DNA microsatellite analysis of multiple-sire matings.

It is important for bovine DNA testing laboratories to provide the cattle industry with accurate estimates of the efficacy and reliability of DNA tests offered so that end users of this technology can adequately assess the cost-benefits of testing. To address these issues for bovine paternity testing, paternity exclusion probability estimates were obtained from breed panel data and were predictive of the efficacy of the DNA tests used in 39 multiple-sire mating groups, involving 5960 calves and 505 bulls. Paternity testing of these mating groups has demonstrated that the majority involve a variable proportion of unknown sires and this impacts on the reliability of sire allocation. Mathematical models based on binomial or beta-binomial probability distributions were used to estimate the reliability of single-sire allocations from multiple-sire matings involving unknown sires. Reliability of 98-99% is achieved when the exclusion probability is 0.99 or greater, after allowing for up to 20% unknown sires. When the exclusion probability drops below 0.90 and there are 20% unknown sires, the reliability is poor, bringing into question the benefits of testing. This highlights the need for DNA testing laboratories to offer paternity tests with an exclusion power of at least 99%.

Animals↗

Inter-rater reliability of Monitor, Senior Monitor and Qualpacs.

This paper describes inter-rater reliability of three quality assessment instruments: Monitor, Senior Monitor and Qualpacs. The work forms part of a Department of Health funded study examining the reliability and validity of these three instruments. Inter-rater reliability is not a fixed entity and therefore should be tested every time raters or the setting change. In this instance, testing was carried out in three wards (one elderly and two surgical). Two methods of analysis were used: percent agreement and intra-class correlation coefficient. Results and some of the techniques that were found to be useful in enhancing reliability are described. Acceptable levels of inter-rater reliability with all three instruments were reached.

Geriatric Nursing↗

Screening for depression in a hepatitis C population: the reliability and validity of the Center for Epidemiologic Studies Depression Scale (CES-D).

RATIONALE: Depression is reported as a serious adverse event of antiviral therapy used to treat patients with hepatitis C (HCV); therefore, there is a need to identify a reliable and valid measure of depressive symptoms for this population. AIMS: To determine reliability, construct validity and predictive validity of the Center for Epidemiological Studies Depression Scale (CES-D) in a hepatitis C (HCV) population. ETHICAL ISSUES: Study reviewed/approved by the University Institutional Review Board and informed consent obtained. METHODS: Longitudinal design testing psychometric properties of the CES-D prior to treatment and 4 and 24 weeks postinitiation of treatment. Reliability was tested using Cronbach's coefficient alpha. Construct validity was tested, prior to therapy, using principal components factoring with varimax rotation. Predictive validity was tested using repeated measures analysis of variance (anova) of CES-D scores at 4 and 24 weeks postinitiation of treatment. RESULTS: Non-probability sample, 116 adult HCV patients [62 (53%) males and 54 (47%) females]. Reliability (Cronbach's alpha) = 0.88 pretreatment, 0.89 week 4 and 0.90 week 24. Construct validity testing revealed four factors: negative affect; positive affect; somatic; and depressed affect/somatic. Exception for two items, 'felt sad' and 'couldn't get going', all items loaded distinctly with correlation coefficients in the range of 0.51-0.84. Predictive validity testing revealed a statistically significant effect over time (P < 0.001) in the direction predicted (pretreatment x = 13.97; post 4 weeks x = 19.54 and 24 weeks x = 19.97). CONCLUSIONS: The CES-D is a reliable and valid instrument to screen for depressive symptoms in HCV patients. The instrument detected the predicted increase in depression associated with HCV. Examination of the sensitivity and specificity is needed to determine the most accurate cut-off score.

Adult↗

Assessment of general practitioners by video observation of communicative and medical performance in daily practice: issues of validity, reliability and feasibility.

OBJECTIVES: To develop a video assessment method for General Practitioners (GPs) by analysing issues of validity, reliability and feasibility of observation of videotaped regular consultations. DESIGN: In a cross-sectional study consultations of 93 GPs were video recorded in the practice during 1 week. The GPs registered consultation and patient data in a logbook; 16 consultations per GP were selected using preset criteria. The quality of communicative and medical performance of these consultations was assessed by GP observers with a validated instrument. The validity of the procedure was evaluated by checking the content of each GP's sample using specific sample criteria. Selection bias was estimated by multiple regression analysis, with sample characteristics as independent variables and scores on communication and medical performance as dependent variables. The influence of observation on GPs and patients was assessed by a questionnaire. Generalizability theory was used to estimate reliability. Feasibility was assessed by conducting a questionnaire, by keeping accounts, and by checking the technical quality of the videotaped consultations. SETTING: Universities of Nijmegen and Maastricht, The Netherlands. SUBJECTS: General Practitioners (GPs). RESULTS: The domain of general practice was well covered in the samples; content validity was satisfactory. With regard to the sample characteristics, only the total duration of consultations appeared to correlate significantly with both the score on communication and the score on medical performance. A majority (71%) of GPs reported not being influenced by the observation, except in the first cases, and recognizing their usual daily performance in the videotaped consultations. An acceptable level of reliability was reached after 2.5 hours of observation, i.e. 12 cases by a single observer. The method was well accepted by both GPs and patients. The costs were pound250 per GP. CONCLUSIONS: Video assessment of GPs in daily practice according to the procedures described is a valid and reliable method, one which is useful for education and quality improvement. There is a trade-off between feasibility on one hand and validity, reliability and credibility on the other hand. Compared to investments in observation methods in standardized settings, the costs of video observation of GPs' actual performance are acceptable.

Clinical Competence↗

The effect on reliability of adding a separate written assessment component to an objective structured clinical examination.

PURPOSE: To determine the effect on test reliability when a separate written assessment component is added to an objective structured clinical examination (OSCE). METHOD: Volunteers (n=38) from Maastricht Medical School were recruited to take a skills-related knowledge test in addition to their regular end-of-year OSCE. The OSCE scores of these volunteers did not differ from those of the other students of their class. Multivariate generalizability theory was used to investigate the combined reliability of the two test formats as well as their respective contributions to overall reliability. RESULTS: Combining the two formats has an added value. The loss of reliability due to the use of fewer stations in the OSCE can be fully compensated by lengthening the written test component. CONCLUSION: From the perspective of test reliability, it is possible to economize on the resources needed for performance-based assessment by adding a separate written test component.

Adult↗

Achieving acceptable reliability in oral examinations: an analysis of the Royal College of General Practitioners membership examination's oral component.

BACKGROUND: The membership examination of the Royal College of General Practitioners (RCGP) uses structured oral examinations to assess candidates' decision making skills and professional values. AIM: To estimate three indices of reliability for these oral examinations. METHODS: In summer 1998, a revised system was introduced for the oral examinations. Candidates took two 20-minute (five topic) oral examinations with two examiner pairs. Areas for oral topics had been identified. Examiners set their own topics in three competency areas (communication, professional values and personal development) and four contexts (patient, teamwork, personal, society). They worked in two pairs (a quartet) to preplan questions on 10 topics. The results were analysed in detail. Generalisability theory was used to estimate three indices of reliability: (A) intercase (B) pass/fail decision and (C) standard error of measurement (SEM). For each index, a benchmark requirement was preset at (A) 0.8 (B) 0.9 and (C) 0.5. RESULTS: There were 896 candidates in total. Of these, 87 candidates (9.7%) failed. Total score variance was attributed to: 41% candidates, 32% oral content, 27% examiners and general error. Reliability coefficients were: (A) intercase 0.65; (B) pass/fail 0.85. The SEM was 0.52 (i.e. precise enough to distinguish within one unit on the rating scale). Extending testing time to four 20-minute oral examinations, each with two examiners, or five orals, each with one examiner, would improve intercase and pass/fail reliabilities to 0.78 and 0.94, respectively. CONCLUSION: Structured oral examinations can achieve reliabilities appropriate to high stakes examinations if sufficient resources are available.

Clinical Competence↗

Reliable microsatellite genotyping of the Eurasian badger (Meles meles) using faecal DNA.

The potential link between badgers and bovine tuberculosis has made it vital to develop accurate techniques to census badgers. Here we investigate the potential of using genetic profiles obtained from faecal DNA as a basis for population size estimation. After trialling several methods we obtained a high amplification success rate (89%) by storing faeces in 70% ethanol and using the guanidine thiocyanate/silica method for extraction. Using 70% ethanol as a storage agent had the advantage of it being an antiseptic. In order to obtain reliable genotypes with fewer amplification reactions than the standard multiple-tubes approach, we devised a comparative approach in which genetic profiles were compared and replication directed at similar, but not identical, genotypes. This modified method achieved a reduction in polymerase chain reactions comparable with the maximum-likelihood model when just using reliability criteria, and was slightly better when using reliability criteria with the additional proviso that alleles must be observed twice to be considered reliable. Our comparative approach would be best suited for studies that include multiple faeces from each individual. We utilized our approach in a well-studied population of badgers from which individuals had been sampled and reliable genotypes obtained. In a study of 53 faeces sampled from three social groups over 10 days, we found that direct enumeration could not be used to estimate population size, but that the application of mark-recapture models has the potential to provide more accurate results.

Alleles↗

Amphetamine withdrawal: I. Reliability, validity and factor structure of a measure.

OBJECTIVE: The aim of this study was to create a short, reliable and valid questionnaire for the evaluation of amphetamine withdrawal, which we shall call the Amphetamine Withdrawal Questionnaire (AWQ). METHOD: Items of the AWQ included in this study were based on the fourth edition of the Diagnostic and Statistical Manual for Mental Disorders (DSM-IV) and a comprehensive review. A field trial for assessing the reliability, validity and factor structure was conducted in outpatients and inpatients with amphetamine withdrawal. RESULTS: Thirty and 102 patients' data were included in the reliability-validity tests and the factor study, respectively. Due to the very low mean score of insomnia item, this item was excluded from subsequent analyses. The AWQ internal consistency was satisfactory with a Cronbach's alpha of 0.77. For test-retest reliability, a Spearman rank order correlation coefficient of the AWQ total score was 0.79. The AWQ total score for criterion validity was moderately correlated with the other two accepted measures. Principal component analysis, eigenvalue-one test and a varimax rotation performed to elicit the factors of AWQ yielded a three-factor model of AWQ: namely hyperarousal, reversed vegetative and anxiety factors. CONCLUSIONS: The AWQ is a short, reliable and valid measure for assessing amphetamine withdrawal symptoms. Further studies with a larger number of patients should be conducted to confirm the results of this factor analysis.

Adolescent↗

Interrater reliability of the Japanese version of the Positive and Negative Syndrome Scale and the appraisal of its training effect.

The purpose of the present study is to test interrater reliability of the Japanese version of the Positive and Negative Syndrome Scale (PANSS) and to examine factors possibly affecting the reliability. The study group conducted the PANSS rating on 20 patients with DSM-IV schizophrenia. For the analysis of interrater reliability, intraclass correlation coefficient (ICC) was calculated. The ICC for individual items of the PANSS ranged from 0.26 to 0.92, and those for the positive, negative, and general psychopathology subscales were 0.85, 0.83 and 0.75, respectively. The Cronbach's alpha coefficient for the subscales were 0.84, 0.87 and 0.76, respectively. The interrater reliability and the internal consistency were satisfactory and similar to those obtained in the antecedent studies. No salient training effect was found in a sequential analysis of the concordance rate. It is concluded that the Japanese version of the PANSS is a reliable and efficient tool for comprehensive assessment of the schizophrenic syndrome.

Adult↗

The Japanese version of the Barratt Impulsiveness Scale, 11th version (BIS-11): its reliability and validity.

No instrument for assessing impulsiveness has been developed in Japan. After translating the Barratt Impulsiveness Scale 11th version (BIS-11) into Japanese, we investigated reliability and validity in student (n = 34) and worker (n = 416) samples. To assess test-retest reliability, the intraclass coefficient between test and retest was calculated in the student sample. Internal consistency was examined by calculating Cronbach's alpha in the worker sample. To see factor validity, we examined by confirmatory factor analysis whether the three-factor model, proposed by a previous report, fit the data. The results showed that the Japanese version of the BIS-11 had excellent test-retest reliability and acceptable internal consistency reliability. In addition, the Japanese version was judged to have similar factor structure to the original one. The Japanese version of the BIS-11 is a reliable and valid measure and has possible utility for assessing impulsiveness.

Adult↗

Reliability of computer-assisted retinal vessel measurementin a population.

The purpose of the study was to assess the intergrader and intragrader reliability of computer-assisted retinal vessel dia-meter measurement in a defined, community-based population. Retinal photographs from participants in the Blue Mountains Eye Study were digitized using standard techniques. A grader identified all retinal vessels located 0.5-1.0 disc diameter from the optic disc margin,and a computer program measured the width of these vessels. Intergrader and intragrader reliability was assessed on a random sub-sample of 184 and 97 images, respectively, using quadratic weighted kappa(kappa) and correlation analysis (R2). Intergrader reliability was high for summary indices of retinal arteriolar (kappa = 0.85, R2 = 0.88)and venular (kappa = 0.90, R2 = 0.90)diameters, and their ratio, the arteriole-to-venule ratio (kappa = 0.75, R2 = 0.79).Intragrader reliability was also high, with kappa values ranging from 0.80 to 0.93 and from 0.80 to 0.92 for graders 1 and 2, respectively. It is concluded that the retinal vessel diameters could be reliably measured using computer-assisted software and may be used for population-based research.

Aged↗