Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

A survey of validity and utility of electronic patient records in a general practice.

OBJECTIVE: To develop methods of measuring the validity and utility of electronic patient records in general practice. DESIGN: A survey of the main functional areas of a practice and use of independent criteria to measure the validity of the practice database. SETTING: A fully computerised general practice in Skipton, north Yorkshire. SUBJECTS: The records of all registered practice patients. MAIN OUTCOME MEASURES: Validity of the main functional areas of the practice clinical system. Measures of the completeness, accuracy, validity, and utility of the morbidity data for 15 clinical diagnoses using recognised diagnostic standards to confirm diagnoses and identify further cases. Development of a method and statistical toolkit to validate clinical databases in general practice. RESULTS: The practice electronic patient records were valid, complete, and accurate for prescribed items (99.7%), consultations (98.1%), laboratory tests (100%), hospital episodes (100%), and childhood immunisations (97%). The morbidity data for 15 clinical diagnoses were complete (mean sensitivity=87%) and accurate (mean positive predictive value=96%). The presence of the Read codes for the 15 diagnoses was strongly indicative of the true presence of those conditions (mean likelihood ratio=3917). New interpretations of descriptive statistics are described that can be used to estimate both the number of true cases that are unrecorded and quantify the benefits of validating a clinical database for coded entries. CONCLUSION: This study has developed a method and toolkit for measuring the validity and utility of general practice electronic patient records.

Databases, Factual↗

RealSpot: software validating results from DNA microarray data analysis with spot images.

The spot images from DNA microarray highly affect the discovery of biological knowledge from gene expression data. However, results from quality analysis, normalization, differential expression, and cluster analysis are rarely validated with spot images in current data analysis methods or software packages. We designed RealSpot, a software package, to validate the results by directly associating spot quality and data with spot images in a spreadsheet table. RealSpot splits hybridization images into individual spots stored in a spreadsheet table. It subsequently associates microarray data with spot images and performs data validation through the standard table operation such as sorting, searching, and editing. RealSpot has several built-in functions to facilitate data validation, including spot quality analysis, data organization, one-way ANOVA, gene ontology association, verification, import, and export. We used RealSpot to evaluate 77 slides (30,000 features each) from real hybridization experiments and to validate results from each step of data analysis. It took approximately 10 min to validate results of spot quality after initial evaluation and correct approximately 0.3% of falsely assigned qualities of 10,000 spots. We validated 1,641 of 2,110 differentially expressed genes identified by SAM analysis in approximately 1/2 h by comparing each gene with its respective spot image. Furthermore, we found that 6 of 48 genes in one cluster from k-mean clustering method showed inconsistent trends of spot images. RealSpot is efficient for validating microarray results and thus helpful for improving the reliability of the whole microarray experiment for experimentalists.

Algorithms↗

Greek M.D. Anderson Symptom Inventory: validation and utility in cancer patients.

OBJECTIVE: The M.D. Anderson Symptom Inventory (MDASI) is a brief assessment of the severity and impact of cancer-related symptoms. The purpose of this study was the translation and validation of the questionnaire in Greek (G-MDASI). METHODS: The translation and validation of the assessment took place at a Pain Relief and Palliative Care Unit. The final validation sample included 150 cancer patients (61 males, 89 females, age range 31-88 years, mean age 63.32). The patients completed the questionnaires at the outpatient clinic. Assessing the validity and reliability constituted the actual validation of the G-MDASI. RESULTS: The item 'diarrhea' had a score of 0 in 139 patients and, thus was omitted from the 'core' list. Consequently, the core questionnaire consisted of 14 items. Factor analysis resulted in a 3-factor model, in both validation and cross-validation samples. The examination of the sensitivity of the MDASI revealed that there were differences between patients with poor-to-good performance status but no differences were found between patients in different treatment groups. CONCLUSIONS: The results showed that the G-MDASI is a reliable and valid measure in Greek cancer patients. It has proved to be a comprehensive symptom assessment tool.

Adult↗

Quality of paediatric rehabilitation from the parent perspective: validation of the short Measure of Processes of Care (MPOC-20) in the Netherlands.

OBJECTIVE: In the present study we aim to assess the reliability and validity of the 20-item version of the Dutch Measure of Processes of Care (MPOC). DESIGN: The reliability, concurrent validity, predictive validity and construct validity of the Dutch MPOC-20 were determined. A subset of MPOC-20 data was extracted from a large Dutch MPOC (56-item version) database. SUBJECTS: Participants were 405 mothers and 22 fathers of children aged 1-18 years recruited through nine paediatric rehabilitation centres in the Netherlands. MAIN MEASURES: The participants filled out the MPOC-20 items, the Client Satisfaction Questionnaire (CSQ), and two additional questions about satisfaction with services and the amount of stress they experienced. RESULTS: The internal consistency analyses (alphas 0.75-0.87) and the test-retest analyses (intraclass correlation coefficients (ICCs) 0.78-0.91) showed that the Dutch MPOC-20 is a reliable tool. The concurrent validity of the Dutch MPOC-20 was confirmed by positive correlations between MPOC-20 scale scores and the CSQ (r 0.39-0.69), and between MPOC-20 scale scores and an overall satisfaction variable (r 0.37-0.66). The predictive validity of the Dutch MPOC-20 was supported by moderately negative correlations between MPOC-20 scores and a stress variable (r -0.27 to -0.44). The construct validity of the Dutch MPOC-20 was confirmed by significant scale intercorrelations (r 0.41-0.84) and a factor analysis. CONCLUSIONS: The 20-item version of the MPOC (Dutch MPOC-20) is a reliable and valid measure of the family-centredness of paediatric rehabilitation.

Adolescent↗

Evaluation of clinical research knowledge and interest among pharmacy residents: survey design and validation.

PURPOSE: The development and validation of a survey to describe the research knowledge, attitudes, and skills of pharmacy practice residents are described. SUMMARY: A survey was drafted to determine if pharmacy practice residency experience and the American Society of Health-System Pharmacists (ASHP)-required project improve the residents' objectively and subjectively assessed research knowledge, to determine if the residency experience and the ASHP-required project affect the residents' attitudes regarding research as a component of their future professional practice, and to subjectively assess the effect of the residency experience and the ASHP-required project on other essential skills, such as problem solving, critical thinking, and time management. An initial questionnaire was developed and underwent content validation testing by clinical pharmacists and faculty, residents, and research fellows. Following content validation, the questionnaire underwent construct validity testing (for discriminative validity and responsiveness) in students, residents, and clinical pharmacists and faculty. Reliability was tested in a subgroup of subjects who completed the questionnaire twice within two to four weeks. From the content validation phase, average scores for individual questions ranged from 1.00 to 2.00. Discriminative validity testing of the revised questionnaire demonstrated the instrument's ability to discriminate between groups expected to differ. Effect-size and mean-knowledge score differences indicated high levels of responsiveness, signifying the instrument's ability to detect change over time or after an intervention. CONCLUSION: A survey questionnaire developed to measure research knowledge and interest among pharmacy practice residents demonstrated its validity and reliability with significant sensitivity and responsiveness.

Data Collection↗

The Acne Quality of Life Index (Acne-QOLI): development and validation of a brief instrument.

BACKGROUND: Acne affects many people and can be detrimental to affected patients' quality of life. Assessing the impact of acne on quality of life requires well-validated and reliable measures of acne-specific quality of life that are brief and easy to administer and interpret. OBJECTIVES: This paper reports on the development and validation of the Acne Quality of Life Index (Acne-QOLI) for use in clinical care, research, and product development. METHODS: Focus groups consisting of people from demographically different populations were conducted to identify the most relevant domains of functioning affected by acne; on the basis of these findings, candidate items were developed. An initial item pool of 58 items was included in a survey of 480 persons with mild to severe acne ranging in age from 12 to 62 years. Factor analysis and qualitative analysis were used to reduce the item pool to 21 items. The construct validity, concurrent validity, internal consistency, and test-retest reliability of the items were evaluated. RESULTS: The 21-item Acne-QOLI showed excellent face validity, content validity, concurrent validity, and construct validity. High internal consistency and test-retest reliability were also found. CONCLUSIONS: Quality of life is now recognized as an important outcome in medical care. The Acne-QOLI is a brief and easily administered and interpreted measure of acne-related quality of life that can be used in clinical care, research, and product development.

Acne Vulgaris↗

[Validity and reliability of the Turkish form of symptom interpretation questionnaire].

OBJECTIVE: This study was undertaken to investigate the validity and reliability of the Turkish version of Symptom Interpretation Questionnaire (SIQ) that was developed by Robbins and Kirmayer (1991). METHOD: Subjects of the study were: 53 patients with medically unexplained somatic complaints (MUS) that were referred to I. Psychiatry Clinic of Ankara Numune Teaching and Research Hospital, 26 patients who were diagnosed as Major Depressive Disorder (MDD) based on their mental complaints, and 71 healthy individuals (Control). RESULTS: Construct validity of SIQ was studied with factor analysis and the scale is statistically shown to consist of three theoretical styles of symptom attribution; normalizing, psychologizing and somatizing. Criterion related validity of normalizing subscale was shown with the absence of significant correlations between SIQ-normalizing scale and either Toronto Alexythimia Scale-20 or SIQ-number of symptoms. Validity of psychologizing subscale was demonstrated by obtaining higher significant mean psychologizing subscale scores in MDD group when compared with mean scores in MUS and Control groups. Criterion related validity of somatizing subscale was shown with significant correlations between SIQ-somatizing scale and SIQ- number of symptoms, and by demonstrating that mean somatizing subscale scores in patients with history of physical disease was significantly higher from mean scores in patients without such history. Internal reliability of SIQ subscales were, alpha=0.86 for normalizing subscale, alpha=0.87 for psychologizing subscale and alpha=0.87 for somatizing subscale. CONCLUSION: Internal reliability, criterion related validity, discriminating power for specific groups and construct validity of the SIQ Turkish version were demonstrated to be satisfactorily valid and reliable.

Adult↗

Validity of an instrument assessing oral health problems in people with Down syndrome.

AIM: The aim of this study was to validate a proxy measure of oral health designed to be completed by the English-speaking parents of people with Down syndrome (DS) aged four years or more. METHODOLOGY: Items were generated through literature review, interviews with parents of people with DS and professional experts and through frequency testing. Data were gathered from one population-based and two clinic-based samples for the separate aspects of validation. Validation consisted of evaluation of: i) internal reliability of the domain structure through Cronbach's alpha; ii) criterion validity against clinical indicators and a clinician's evaluation of some items; iii) construct validity involving an age-matched comparison of domain scores between people with DS and non-DS siblings, and within the DS group by health status indicators; and iv) test-retest reliability through the generation of intra-class correlation coefficients (ICC). RESULTS: A 20-item instrument with four domains (communication, eating, parafunction and symptoms) was developed. Cronbach's alpha by domain was 0.5-0.8. Indicators of criterion validity for domains against clinical indicators (Spearman's coefficient 0.1-0.4) and parent-rated items against clinician-rated items (weighted Kappa 0.1-0.8) were varied as anticipated. Indicators of construct validity (differences with non-DS siblings and correlations with medical status within the DS group) were excellent. Test-retest reliability was good (ICC range 0.64-0.84). CONCLUSION: These data suggest the test instrument is valid as a descriptive, discriminative, proxy English language measure of oral health problems in people with DS aged four years or more.

Adolescent↗

Reliability and construct validity of the Malay version of the Job Content Questionnaire (JCQ).

The JCQ has been shown to be a valid and reliable instrument to assess job stress in many occupational settings worldwide. In Malaysia, both the English and validated Malay versions have been employed in studies of medical professionals and laboratory technicians, respectively. The present study assessed the reliability and construct validity of the Malay version of the JCQ among automotive workers in Malaysia. Fifty workers of a major automotive manufacturer in Kota Bharu, Kelantan consented to participate in the study and were administered the Malay version of the JCQ. Translation (English-Malay) and back translation (Malay-English) of the JCQ was made to ensure the face validity of the questionnaire. Reliability was determined using Cronbach's alpha for internal consistency, whilst construct validity was assessed using exploratory factor analysis (principal component with varimax rotation). The results indicate that the Cronbach's alpha coefficients were acceptable for decision latitude (job control or decision authority) (0.74) and social support (0.79); however, it was slightly lower for psychological job demand (0.61). Exploratory factor analysis showed 3 meaningful common factors that could explain the 3 theoretical dimensions or constructs of Karasek's demand-control-social support model. In conclusion, the results of the validation study suggested that the JCQ scales are reliable and valid for assessing job stress in a population working in the automotive industry. Further analyses are necessary to evaluate the stability and concurrent validity of the JCQ.

Adult↗

Reliability and validity of refractive error-specific quality-of-life instruments.

OBJECTIVE: To evaluate the reliability and validity of the National Eye Institute Refractive Error Quality of Life Instrument (NEI-RQL-42) and the Refractive Status and Vision Profile survey (RSVP). METHODS: Eighty-one participants with good visual acuity (better than 20/30 best-corrected acuity in each eye) completed the NEI-RQL-42 and RSVP on 2 occasions. Noncycloplegic, subjective refractions and high-contrast visual acuity assessments were also performed. Statistical analyses addressed internal consistency, test-retest reliability, and validity (ie, concurrent and construct validity) of the 2 instruments. OUTCOME MEASURES: The NEI-RQL-42, RSVP survey, subjective refraction, and visual acuity. RESULTS: The internal consistency for the overall NEI-RQL-42 was excellent (Cronbach alpha = 0.91); and for the overall RSVP, good (Cronbach alpha = 0.81). Likewise, the test-retest reliability for the overall NEI-RQL-42 was excellent (intraclass correlation coefficient [ICC], 0.91; 95% limits of agreement, -9.1 to 10.1); and for the RSVP, fair (ICC, 0.76; 95% limits of agreement, -12.1 to 12.5). The NEI-RQL-42 overall score showed good concurrent validity as it correlated significantly with subjective refraction, whereas the RSVP overall score did not. The NEI-RQL-42 and RSVP showed similar construct validity in terms of refractive error discrimination, but the NEI-RQL-42 showed better construct validity when discriminating by the type of refractive correction used by patients. Between-instrument convergent and divergent validity was good. CONCLUSIONS: The NEI-RQL-42 and RSVP generally have good reliability and validity in this sample of patients with refractive error. However, other factors such as content should be considered in choosing 1 of these instruments for studies of refractive error correction.

Adult↗

Development and validation of the effectiveness of [corrected] auditory rehabilitation scale.

OBJECTIVE: To develop a new scale of hearing-related function and quality of life in patients with hearing aids that addresses overlooked concerns, such as hearing-aid comfort, convenience, and cosmetic appearance, that may influence hearing-aid adherence while maintaining brevity and sensitivity to clinical change. DESIGN: Prospective, multicenter instrument validation. SETTING: Four diverse sites in Washington State, including 2 private practices, 1 university setting, and 1 Veterans Affairs hospital. PATIENTS: Seventy-eight patients with hearing aids. INTERVENTIONS: We created 2 modules in the Effectiveness of Auditory Rehabilitation (EAR) scale. The first module (Inner EAR) covers intrinsic hearing issues such as hearing in quiet and hearing in noise and is administered both before and after treatment. The second module (Outer EAR) covers extrinsic (hearing-aid related) issues such as comfort, appearance, and convenience and is administered after hearing-aid fitting. MAIN OUTCOME MEASURES: Both scales were developed and validated in 3 stages. Stage 1 used a qualitative approach from multiple data sources to develop preliminary instruments. Stage 2 used approaches from classic test theory to reduce the number of items and psychometrically validate the instruments. Stage 3 examined the responsiveness or sensitivity to clinical change. RESULTS: A 10-item Inner EAR module and a 10-item Outer EAR module were created and validated. Internal consistency of individual domains (Cronbach alpha = 0.85 and 0.72, respectively) and test-retest reliability (intraclass correlation coefficients = 0.76 and 0.81, respectively) were excellent. Evidence of construct validity included concurrent validity with other hearing scales and global visual analog scales, discriminant validity with dizziness handicap, correlation with hearing-aid adherence, and confirmatory factor analyses. Both scales had strong evidence of responsiveness (sensitivity to change), with higher effect sizes and Guyatt responsiveness statistics than the 2 widely used hearing scales in this study. The scales took an average of 5 minutes to complete. CONCLUSIONS: The EAR scale is a valid and reliable measure of the effectiveness of amplification in the treatment of sensorineural hearing loss. It addresses the range of issues that are of importance to hearing-aid patients. The scales have excellent psychometric properties, are more responsive than several widely used hearing scales, and are minimally burdensome for patients to complete. The EAR may be a valuable outcome measure in future studies of both existing hearing aids and newer hearing-aid technologies, such as bone-anchored aids or middle ear implants.

Factor Analysis, Statistical↗

Efficient regression calibration for logistic regression in main study/internal validation study designs with an imperfect reference instrument.

An extension to the version of the regression calibration estimator proposed by Rosner et al. for logistic and other generalized linear regression models is given for main study/internal validation study designs. This estimator combines the information about the parameter of interest contained in the internal validation study with Rosner et al.'s regression calibration estimate, using a generalized inverse-variance weighted average. It is shown that the validation study selection model can be ignored as long as this model is jointly independent of the outcome and the incompletely observed covariates, conditional, at most, upon the surrogates and other completely observed covariates. In an extensive simulation study designed to follow a complex, multivariate setting in nutritional epidemiology, it is shown that with validation study sizes of 340 or more, this estimator appears to be asymptotically optimal in the sense that it is nearly unbiased and nearly as efficient as a properly specified maximum likelihood estimator. A modification to the regression calibration variance estimator which replaces the standard uncorrected logistic regression coefficient variance with the sandwich estimator to account for the possible misspecification of the logistic regression fit to the surrogate covariates in the main study, was also studied in this same simulation experiment. In this study, the alternative variance formula yielded results virtually identical to the original formula. A version of the proposed estimator is also derived for the case where the reference instrument, available only in the validation study, is imperfect but unbiased at the individual level and contains error that is uncorrelated with other covariates and with error in the surrogate instrument. Replicate measures are obtained in a subset of study participants. In this case it is shown that the validation study selection model can be ignored when sampling into the validation study depends, at most, only upon perfectly measured covariates. Two data sets, a study of fever in relation to occupational exposure to antineoplastics among hospital pharmacists and a study of breast cancer incidence in relation to dietary intakes of alcohol and vitamin A, adjusted for total energy intake, from the Nurses' Health Study, were analysed using these new methods. In these data, because the validation studies contained less than 200 observations and the events of interest were relatively rare, as is typical, the potential improvements offered by this new estimator were not apparent.

Adult↗

Predictive, concurrent, prospective and retrospective validity of self-reported delinquency.

BACKGROUND: The self-report method is widely used to measure offending. Previous studies suggest that it is generally valid, but that its validity may be lower for blacks than for whites. AIM To assess the validity of self-reported offending in relation to court referrals, and to investigate how it varies with types of offences, sex and race. METHOD: Annual court and self-report data were collected between ages 11 and 17 for eight offences in the Seattle Social Development Project, which is a prospective longitudinal survey of 808 youths. RESULTS: Self-reports predicted future court referrals. Predictive validity was highest for drug offences, for males and for whites, and lowest for females and Asians. The probability of youths with a court referral reporting offences and arrests was highest for drug offences, for males, for whites and for blacks. Retrospective ages of onset agreed best with prospective ages for drug offences, Asians and whites. More Asians than blacks or whites failed retrospectively to report offences that had been reported prospectively. CONCLUSIONS: The validity of self-reports of offending was high, especially for drug offences, for males and for whites. Contrary to prior research, validity was high for black males. It was lowest for Asian females. Sex and race differences in validity held up after controlling for socioeconomic status. Differential validity probability did not reflect police bias.

Adolescent↗

QUALIDEM: development and evaluation of a dementia specific quality of life instrument--validation.

OBJECTIVE: To validate the QUALIDEM, a quality of life measure for people with dementia within residential settings rated by professional caregivers. METHOD: In a sample of 202 residents of nursing homes Spearman rank correlations were calculated between the QUALIDEM subscales aand indices of convergent validity and discriminant validity, with dementia severity and need of care, with global QOL scores by the head nurse and family, and with self-report on COOP/WONCA Charts. RESULTS: The one-method multi-trait matrix showed 90.5% of the correlations to be in support for convergent and discriminant validity. Low to moderate correlations were observed with dementia severity and need of care, confirming that QOL is not merely disease severity. Support for concurrent validity was found in correlations with QOL ratings by the head nurse. The QUALIDEM did not correlate with most of the family ratings or with the COOP/WONCA Charts. CONCLUSION: The results of this validation study together with the obtained content validity through the method of construction provide sufficient support for validity of the QUALIDEM to be used for care evaluation and research in residential settings.

Activities of Daily Living↗

Does the presence of anxiety affect the validity of a screening test for depression in the elderly?

INTRODUCTION: Depression in the elderly is frequently detected by screening instruments and often accompanied by anxiety. We set out to study if anxiety will affect the ability to detect depression by a screening instrument. OBJECTIVE: To validate the short Zung depression rating scale in Israeli elderly and to study the affect of anxiety on its validity. DESIGN: The short Zung was validated against a psychiatric evaluation, in a geriatric inpatient and outpatient service. The overall validity was determined, as well as for subgroups of sufferers and non-sufferers of anxiety. SETTING: An urban geriatric service in Israel. PATIENTS: 150 medical inpatients and outpatients, aged 70 years and older. MEASURES: Psychiatric evaluation of modified Anxiety Disorders Interview Schedule for DSM-IV as criterion standard for anxiety and depression and short Zung instrument for depression. RESULTS: By criterion validity, 60% suffered from depression. The overall validity of the short Zung was high (sensitivity 71.1%, specificity 88.3%, PPV 90.1%, NPV 67.1%). The validity for those not suffering from anxiety was good (sensitivity 71.1%, specificity 90.2%, PPV 84.4%, NPV 80.7%). In those with anxiety, sensitivity, specificity and PPV were high (71.2%, 77.8%, 94.9% respectively), although the specificity was less than in non-suffers. However major difference was in the NPV rate being much lower (31.8%). CONCLUSION: The short Zung, an easily administered instrument for detecting depression, is also valid in the Israeli elderly. However, anxiety limits the usefulness of this instrument in correctly ruling out depression. The clinician must be aware, therefore, that those suffering from anxiety may score negatively for depression on a screening instrument, such as the short Zung.

Aged↗

Development and validation of a clinical scale for the diagnosis of drug-induced hepatitis.

The objective of this study is to present and validate a clinical scale for the diagnosis of drug-induced liver injury (DILI). Five components were selected to be included in the scale: temporal relationship between drug intake and the onset of clinical picture, exclusion of alternative causes, extrahepatic manifestations, rechallenge or accidental re-exposure, and previous report in medical literature. The relative importance of each component was weighed, and arbitrary scores were attributed. The probability of the diagnosis of DILI was expressed as a final score, which could vary from -6 to 20. Content validity, criterion validity, construct validity, and inter-rater reliability were studied. To analyze validity and reliability, a random sample of 50 cases of suspected DILI was drawn from a series of 120 cases reported to our unit. The classification of the 50 cases by three experts in DILI was used as the external standard in the study of criterion validity. Agreement between the scale and the standard, and agreement between two independent raters (inter-rater reliability) was analyzed by weighted kappa coefficient. There was agreement between the scale and the standard in 42 cases (84%) with a weighted kappa coefficient of 0.90. A good discriminatory capacity of the scale was found when construct validity was studied. Agreement between raters was observed in 86% of the cases, corresponding to the weighted kappa of 0.93. In conclusion, the clinical scale was shown to have a high-level of validity and inter-rater reliability as well as a good discriminatory capacity between different levels of probability. These data suggest that the scale is suitable for use in clinical practice and may contribute to overcome the difficulties in the process of causality assessment in DILI.

Bias↗

Reliability and validity of clinical outcome measurements of osteoarthritis of the hip and knee--a review of the literature.

High reliability and validity of clinical rating schemes is crucial for their use as outcome measurements of treatment of hip and knee osteoarthritis. In this paper, we review the empirical evidence on the reliability and validity of commonly used clinical scores. Clinical scores and related reliability and validity studies were identified by systematic literature search. Scores were classified according to the type and joint. Reliability and validity studies were characterized according to design, population, number and qualification of observers, number of measurements, time interval between repeat measurements and results. Reliability and validity studies were reported for only 6 and 15 of the 45 identified clinical scores, respectively. Although comparisons are difficult due to differences in study design, relatively high reliability was reported for most measurements of pain, stiffness, and physical function, while results are less conclusive for clinical signs. Most validity studies focused on the correlation between various scores. Correlation was generally found to be high for overall numerical ratings, but scores often differed with respect to the interpretation of these ratings. Validity has been more comprehensively studied for Lequesne's scores, WOMAC, and ILAS, and these scores have shown satisfactory responsiveness to different treatment effects. Overall, knowledge on reliability and validity of clinical scores of hip and knee osteoarthritis is limited, underlining the need for further properly designed and conducted studies.

Hip↗

A general technique for automatic left ventricle boundary validation: relation between gray scale cardioangiograms and observed boundary errors.

This article presents an automatic left ventricle boundary validation technique using the gray scale cardioangiograms, the observed boundary errors, and the left ventricle boundaries from any source that needs to be validated. This validation technique is based on the gray scale information near the boundary of the left ventricle in the cardioangiograms. Using a mutually exclusive window of fixed size, which is centered on the left ventricle boundary vertex and along the left ventricle contour, we compute a difference in contrast value for areas of the window both inside and outside the left ventricle region. These contrast values then are regressed against the observed boundary errors. The observed boundary errors are computed using the polyline distance measure, by comparing two sets of boundaries: boundaries estimated from any boundary estimation algorithm, and the original ground truth boundaries as traced by the cardiologist. We performed our experiments on a database of 245 patient studies, each having two frames: end-diastole (ED) and end-systole (ES). The mean boundary error before running the validation system was 4.4 mm. Using our boundary validation system, by rejecting 5% to 10% of the patient studies, the validation system results in an error of 4.0 mm for the cross-validation case and 3.85 mm for the ideal case. We show the reliability curves of our validation system by computing the probability of false alarm, probability of mis-detection, and mean predicted errors when a total of n patients are rejected from the database of 245 studies.

Algorithms↗