Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Reliability and validity of a new scale to assess postoperative dysfunction after resection of upper gastrointestinal carcinoma.

PURPOSE: We evaluated the purpose reliability and validity of a preliminary scale, which we developed to assess postoperative dysfunction after surgery for gastric and esophageal carcinoma. METHODS: After interviews with 12 patients, reviews of previous studies, and discussions with experts, we identified the physical symptoms that develop after resection of upper gastrointestinal (GIT) carcinoma, and devised a preliminary scale comprised of 34 items. A questionnaire survey based on this scale was then sent to 283 patients. RESULTS: The questionnaire was returned by 223 patients (78.8%), and 219 responses (98.2%) were valid. Among the 219 respondents, 168 had gastric carcinoma and 51 had esophageal carcinoma. After the elimination of scale items regarded as irrelevant based on statistical considerations and the judgment of experts, factor analysis was done. Seven factors were valid, namely, limited activity due to decreased food consumption, reflux, gastric dumping, nausea and vomiting, deglutition difficulty, pain, and difficulty with passing stools, which were often poorly formed. Scale reliability was confirmed by a Cronbach alpha-coefficient of 0.924. The validity of the construction of this scale was confirmed using the known-group technique based on the operative procedures performed, and the results of factorial validity. CONCLUSION: Our preliminary scale is sufficiently reliable and valid, and will prove to be clinically useful.

Adult↗

The validity of children's self-reported exposure to traffic.

This paper reports on a series of studies of the validity of children's self-reported exposure to traffic. The studies were conducted within a larger case-control study carried out in the metropolitan area of Perth, Western Australia. For these validity studies, subjects were randomly selected from the original pool of case (n = 100) and control (n = 400) subjects. Three techniques were used to assess the validity of the self-reported 'habitual exposure', namely, the 'moving observer' technique, pedestrian diaries, and a test of construct validity. The child's regular walking activities during the course of a typical week was considered the child's 'habitual exposure'. The 'moving observer' technique involved either the researcher or research assistant following a random sample of children on their walking trips. A further sample of children maintained a 'diary' of their weekly walking trips, and a number of constructs, in this case variables which related, indirectly, to the child's level of exposure to the road environment, were included in the interview schedule. As two researchers were involved in all aspects of this study, intra- and inter-rater reliability were assessed and tape-recordings of the interviews were used to determine the reliability of the coded data. No significant differences were found between the child's reported exposure to the road environment and either the observed exposure or exposure recorded in pedestrian diaries. For some exposure variables, namely, the number and duration of walking trips, children tended to underestimate their exposure (when compared to the observations of the researchers). For case subjects, the number of roads crossed was also underestimated. The test of construct validity indicated that time spent in 'out-of-home activities' (activities other than going to or from school, which involved time spent in the road environment) does not correlate strongly with 'habitual exposure'. Intra- and inter-rater reliability and the reliability of transcribed data from taped interviews were all found to be very sound. The studies indicate that children's self-reported 'habitual exposure' data is a valid measure of his or her actual exposure in the road environment.

Case-Control Studies↗

Concurrent validity of two language screening tests.

The importance of ascertaining the validity of clinical instruments used to make decisions about individuals is discussed and the need for additional validation studies is emphasized. Steps that can be taken to confirm the validity for a particular application, setting, or population are described. As an example, the concurrent validity of two language screening instruments, the Fluharty Preschool Screening Test and the Northwestern Syntax Screening Test, and their subtests was examined. Decisions from these screening tests and subtests were compared to a validity criterion of passing or failing the Sequenced Inventory of Communication Development for 182 white middle-class children, ages 36-47 months. The results showed that the screening tests differed in their validity, depending upon the content of the test and each subtest. The consequences of using either screening test are explored, to illustrate how the outcomes of such studies should be interpreted.

Child, Preschool↗

Validity criteria for exposure assessment methods.

The validity of any exposure assessment is an essential step in risk assessment. Following definitions of the various types of measurement validity, such as content, criterion and construct validity, and study validity, some shortcomings of actual exposure measurements are discussed. Both direct and indirect ways of measuring exposures to air, water and food pollutants are taken into consideration, with special attention being payed to indirect measurements using questionnaires and/or time-activity budgets. It is concluded that it is difficult to measure validity in an absolute way. Rather, a relative evaluation is possible comparing the validity of one result or study to that of another.

Environmental Exposure↗

Development and validation of an instrument to measure satisfaction of participants at breast screening programmes.

A reliable and valid questionnaire has been developed to measure the satisfaction of participants with service offered at mammography screening programmes. The questionnaire measures five specific aspects: convenience and accessibility, staffs' interpersonal skills, information transfer between staff and client, physical surroundings and perceived technical competence of staff. A general satisfaction dimension was also included. Systematic procedures were followed to ensure that the initial pool of items met the criteria for satisfactory content validity. These procedures included extensive literature review and interviews with participants and service providers. Discriminant validity was assessed by a modified Q-sort procedure, where eight expert judges sorted items into relevant dimensions. The sample for other validity and reliability testing consisted of 584 women who were participants at a breast X-ray programme in Melbourne, Australia. Concurrent validity was demonstrated by considering the correlation of the sum of the subscale scores for each respondent with their score on the general subscale (r = 0.76; P less than 0.001). Multiple regression was used to provide further evidence for the discriminant validity of the proposed subscales and support for the multidimensional conceptualism of satisfaction. Scores on the general satisfaction subscale were used as an outcome variable and other subscale scores were predictor variables. All subscale scores significantly contributed to the prediction of satisfaction, over and above that of other subscales (R2 = 0.59). This indicates that these subscales are measuring distinct dimensions of satisfaction. Cronbach's alpha of each subscale was over 0.50, indicating that the subscales are reliable. The instrument is a potentially useful tool for assessing the quality of care at mammographic screening services and could be used routinely by such services to monitor satisfaction.

Australia↗

Comparative study of the validity of four French McGill Pain Questionnaire (MPQ) versions.

Four different French versions of the McGill Pain Questionnaire (MPQ) have been published: 3 are MPQ translations in Canadian French and 1 (QDSA) is an MPQ reconstruction in (France) French. The aim of our work was to study the validity of these available questionnaires for use in France. The validity was evaluated by 44 French physicians. Various validity criteria were studied: item, dimension, subclass and pain descriptor intensity. A new French MPQ was also developed. Significant validity differences emerged between the different MPQ versions. This study confirms the satisfactory validity of the QDSA. The validity of the newly developed French MPQ was equal but not better than the QDSA. A 15-item short MPQ-QDSA version was also developed. For studies with patients from France, it is recommended that the QDSA or the short MPQ-QDSA versions be used.

Evaluation Studies as Topic↗

Biochemical validation of smoking status: pros, cons, and data from four low-intensity intervention trials.

Biochemical validation of smoking status has long been considered essential, but recent reports have questioned its utility in certain kinds of field trials. We describe efforts to biochemically validate self-reports of smoking cessation from participants in four large-scale randomized trials in outpatient clinics, hospitals, worksites, and dental clinics. These studies included over 5,000 adults smokers who participated in the population-based low-intensity intervention evaluations. At a 1-year follow-up, 798 subjects reported no tobacco use. We attempted to verify these reports using saliva continine/carbon monoxide validation procedures. Overall, there was a moderately high nonparticipation rate (27%), a low disconfirmation rate (4%), and a high self-reported relapse rate (12%) in the interval between survey and biochemical validation. There were no differences between intervention and control conditions on any of the above variables. Longer durations of self-reported abstinence were strongly related to increased probability of biochemical confirmation. Differences in results across projects were related to how biochemical validation was conducted. These results, as well as statistical power considerations, raise questions about whether biochemical validation procedures are practical, informative, or cost-effective in such population-based, low-intensity intervention research.

Adult↗

Development and validation of a logistic regression-derived algorithm for estimating the incremental probability of coronary artery disease before and after exercise testing.

OBJECTIVES: Our goals were to develop and validate a multivariate algorithm for estimating the incremental probability of the presence of coronary artery disease. BACKGROUND: Multivariate methods, including logistic regression analysis, have been extensively applied to diagnostic exercise testing. However, few previous studies have included both an incremental design and external validation. METHODS: A retrospective collection of clinical, exercise test and catheterization data was performed involving four U.S. referral medical centers. All patients had no prior history of coronary disease and had undergone coronary angiography < or = 3 months after exercise stress testing. An algorithm was developed in one center (590 patients with a 41% prevalence of coronary artery disease) with the use of logistic regression analysis and was validated in the other three centers (1,234 patients, 70% prevalence). The algorithm incorporated pretest variables (age, gender, symptoms, diabetes, cholesterol), exercise electrocardiographic (ECG) variables (mm of ST segment depression, ST slope, peak heart rate, metabolic equivalents [METs], exercise angina) and one thallium variable. Discrimination was measured with receiver operating characteristic curve analysis. Calibration (that is, reliability) was assessed from a comparison of probability estimates and the actual prevalence of disease. RESULTS: The overall incremental receiver operating characteristic curve areas for the validation group were pretest, -0.738 +/- 0.016; postexercise ECG, 0.78 (SE 0.017); and postthallium, 0.82 (SE 0.016); p < 0.01 for both increments. Within the three validation institutions, the institution with a disease prevalence closest to that of the derivation institution had the best incremental receiver operating characteristic curve areas. There was a stepwise incremental improvement in calibration especially from exercise ECG to thallium testing. CONCLUSIONS: An incremental multivariate algorithm derived in one center reliably estimated disease probability in patients from three other centers. The incremental value of testing was best demonstrated when the derivation and validation groups had a similar disease prevalence. This algorithm may be useful in decision making that relates to the diagnosis of coronary disease.

Algorithms↗

Symptom validity assessment: practice issues and medical necessity NAN policy & planning committee.

Symptom exaggeration or fabrication occurs in a sizeable minority of neuropsychological examinees, with greater prevalence in forensic contexts. Adequate assessment of response validity is essential in order to maximize confidence in the results of neurocognitive and personality measures and in the diagnoses and recommendations that are based on the results. Symptom validity assessment may include specific tests, indices, and observations. The manner in which symptom validity is assessed may vary depending on context but must include a thorough examination of cultural factors. Assessment of response validity, as a component of a medically necessary evaluation, is medically necessary. When determined by the neuropsychologist to be necessary for the assessment of response validity, administration of specific symptom validity tests are also medically necessary.

Diagnosis, Differential↗

Current concepts in validity and reliability for psychometric instruments: theory and application.

Validity and reliability relate to the interpretation of scores from psychometric instruments (eg, symptom scales, questionnaires, education tests, and observer ratings) used in clinical practice, research, education, and administration. Emerging paradigms replace prior distinctions of face, content, and criterion validity with the unitary concept "construct validity," the degree to which a score can be interpreted as representing the intended underlying construct. Evidence to support the validity argument is collected from 5 sources: CONTENT: Do instrument items completely represent the construct? RESPONSE PROCESS: The relationship between the intended construct and the thought processes of subjects or observers. INTERNAL STRUCTURE: Acceptable reliability and factor structure. RELATIONS TO OTHER VARIABLES: Correlation with scores from another instrument assessing the same construct. CONSEQUENCES: Do scores really make a difference? Evidence should be sought from a variety of sources to support a given interpretation. Reliable scores are necessary, but not sufficient, for valid interpretation. Increased attention to the systematic collection of validity evidence for scores from psychometric instruments will improve assessments in research, patient care, and education.

Humans↗

Typologies of drug dependence: comparative validity of a multivariate and four univariate models.

Data from a longitudinal cohort study were used to directly compare the concurrent and predictive validity of four univariate typologic approaches with a multivariate approach in subtyping drug dependence. The four univariate typologies were based upon: (a) age-of-onset of drug abuse/dependence, (b) presence of drug abuse in first-degree relatives, (c) presence of antisocial personality disorder, and (d) sex. The multivariate typologic approach was based on indices of vulnerability, chronicity, consequences, and psychopathology, yielding the Type A/B dichotomy first demonstrated in alcohol dependence. Subtypes generated from the univariate typologies were then each compared with the multivariate typology on measures of concurrent and predictive validity, and the strength of association was compared statistically. There was evidence of significantly greater concurrent validity of the Type A/B typology compared with the univariate typologies across all the domains of validation (risk, substance use, psychopathology, personality, and overall functioning). The multivariate typology also fared better than the univariate ones in all three domains on which predictive validity was evaluated: substance use, psychopathology, and overall functioning, as well as the degree of change in several composite scores (drug, medical, legal, and psychiatric) and the global psychiatric symptom index. This direct method of comparison seemed to demonstrate the superior validity of the multivariate cluster-analytic approach over the univariate approaches to classifying subjects with drug dependence.

Age Factors↗

DSM-IV alcohol dependence and abuse: further evidence of validity in the general population.

BACKGROUND: In order to understand the validity of the Diagnostic and Statistical Manual of Mental Disorders, 4th ed. (DSM-IV) alcohol abuse and dependence diagnoses, studies are needed in both clinical and general population samples. The purpose of this study was to examine the construct and criterion-oriented validity of DSM-IV alcohol dependence and abuse in the general population with respect to factor structure and their relationship to family history of alcoholism, treatment utilization, and psychiatric comorbidity. METHODS: This analysis is based on data from the 2001-2002 National Epidemiologic Survey on Alcohol and Related Conditions (NESARC), in which nationally representative data were collected in personal interviews conducted with one randomly selected adult in each sample household or group quarters. A subset (n=26,946) of the NESARC sample (total n=43,093) who reported drinking one or more drinks during the year preceding the interview formed the basis of analyses. Latent variable modeling was used to assess the concurrent validity of DSM-IV alcohol abuse and dependence symptom items. RESULTS: The latent variable modeling yielded one major factor related to alcohol dependence, a second factor related to alcohol abuse and a third smaller factor defined by tolerance. The validity of alcohol dependence in general population samples was further supported by statistically significant associations with family history of alcoholism, treatment utilization, and psychiatric and medical comorbidities. CONCLUSIONS: The factor structure and relationship to external criterion variables observed in the study provide support for the further validity of DSM-IV alcohol dependence in the general population, whereas support for the validity of DSM-IV abuse was equivocal.

Adolescent↗

Teachers' ratings of gross motor skills suffer from low concurrent validity.

In this study an attempt was made to construct a reliable and valid unifactorial teachers' rating scale for gross motor ability. Study 1 (132 children from 3 to 7 years) revealed that reliability of the scale was acceptable and that the scale represented an unifactorial dimension. Two studies on concurrent validity of the scale with an experimental gross motor task (stepping-stone crossing), the unifactorial subtest Locomotion of the Test of Gross Motor Development and the subtest Balance of the Movement Assessment Battery for Children as criterion measures, did not produce acceptable validity coefficients. In both validation studies an age effect was found. It was concluded that factor specificity does not seem the answer to the usual low validity coefficients of multifactorial teachers' rating scales. An alternative approach is suggested in which the assessment of functional activities in daily situations is stressed. Finally, the inclusion of atypical groups in random samples, which is common practice in research on concurrent validity of screening instruments for children's motor problems, is discussed.

Child↗

The diagnosis of heart failure in the community. Comparative validation of four sets of criteria in unselected older adults: the ICARe Dicomano Study.

OBJECTIVES: We sought to compare construct and predictive validity of four sets of heart failure (HF) diagnostic criteria in an epidemiologic setting. BACKGROUND: The prevalence estimates of HF vary broadly depending on the diagnostic criteria. METHODS: Data were collected in a survey of community dwellers who were > or =65 years of age living in Dicomano, Italy. At baseline, HF was diagnosed with the criteria of the Framingham, Boston, and Gothenburg studies and of the European Society of Cardiology (ESC). Left ventricular mass index and ejection fraction, left atrium systolic dimension, lower extremity mobility disability, summary physical performance score, and 6-min walk test were compared between HF and non-HF participants to test for construct validity of each set of criteria. Predictive validity was evaluated with follow-up assessment of cardiovascular mortality, incident disability, and HF-related hospitalizations. Comparisons were adjusted for demographics, comorbidity, and psychoaffective status. RESULTS: Of 553 participants, 11.9%, 10.7%, 20.8%, and 9.0% had HF, according to Framingham, Boston, Gothenburg, and ESC criteria, respectively. In terms of construct validity, Framingham and Boston criteria discriminated HF from non-HF participants better than Gothenburg and ESC criteria across the measures of cardiac function and global performance. The Boston criteria showed a superior predictive validity because they indicated a significantly greater adjusted risk of cardiovascular death (hazard ratio3.9, 95% confidence interval 1.2 to 13.2), incident disability, and hospitalizations in participants with HF. CONCLUSIONS: The Boston criteria are preferable to Framingham, Gothenburg, and ESC criteria for the diagnosis of HF in older community dwellers because they have good construct validity and more accurately predict cardiovascular death, incident disability, and hospitalizations.

Aged↗

Development and validation of challenge materials for double-blind, placebo-controlled food challenges in children.

BACKGROUND: The use of double-blind, placebo-controlled food challenges (DBPCFCs) is considered the gold standard for the diagnosis of food allergy. Despite this, materials and methods used in DBPCFCs have not been standardized. OBJECTIVE: The purpose of this study was to develop and validate recipes for use in DBPCFCs in children by using allergenic foods, preferably in their usual edible form. METHODS: Recipes containing milk, soy, cooked egg, raw whole egg, peanut, hazelnut, and wheat were developed. For each food, placebo and active test food recipes were developed that met the requirements of acceptable taste, allowance of a challenge dose high enough to elicit reactions in an acceptable volume, optimal matrix ingredients, and good matching of sensory properties of placebo and active test food recipes. Validation was conducted on the basis of sensory tests for difference by using the triangle test and the paired comparison test. Recipes were first tested by volunteers from the hospital staff and subsequently by a professional panel of food tasters in a food laboratory designed for sensory testing. Recipes were considered to be validated if no statistically significant differences were found. RESULTS: Twenty-seven recipes were developed and found to be valid by the volunteer panel. Of these 27 recipes, 17 could be validated by the professional panel. CONCLUSION: Sensory testing with appropriate statistical analysis allows for objective validation of challenge materials. We recommend the use of professional tasters in the setting of a food laboratory for best results.

Allergens↗

Translation, cross-cultural adaptation, reliability, and validity of the German version of the Coping Strategies Questionnaire (CSQ-D).

UNLABELLED: The aim of this study was to translate and cross-culturally adapt the American version of the Coping Strategies Questionnaire (CSQ) and to test the reliability and validity of the German version (CSQ-D). The CSQ was translated and cross-culturally adapted following international guidelines. Reliability and validity were tested in 62 individuals with chronic musculoskeletal pain syndromes. For the concurrent criterion-related validity the CSQ-D scales were compared with the German Pain Coping Questionnaire (FESV-BW), and for the construct validity with the German Short Form 36 (SF-36). The translation process proceeded without major difficulties. In testing for reliability, the CSQ-D as a whole had a Cronbach's alpha of .94 and an intraclass correlation coefficient of .89 (95% CI .86-.98). The total CSQ-D score was correlated to the FESV-BW scales with scores of r = 0.32-0.55 and with the SF-36 Mental Component Summary with scores of r = 0.32-0.53. The CSQ-D is a precisely translated and highly reliable instrument in the assessment of chronic pain coping strategies. Its concurrent criterion-related validity and construct validity are low. The main reason for the low level of agreement between the CSQ-D and the FESV-BW was revealed by factor analysis. PERSPECTIVE: This paper presents the German version of the Coping Strategies Questionnaire (CSQ-D) together with the results of clinimetric testing. The CSQ-D is a feasible and reliable outcome measure to be used in trials with German-speaking patients or large multicenter multinational trials to assess pain coping strategies in patients with chronic musculoskeletal pain.

Adaptation, Psychological↗

Reliability, validity, and responsiveness of the simple shoulder test: psychometric properties by age and injury type.

The purpose of this study was to measure the reliability, validity, and responsiveness of the Simple Shoulder Test (SST) scale and to examine these in patients stratified by age and injury type. Test-retest reliability, content validity, criterion validity, construct validity, and responsiveness were determined for the SST. The study population comprised 1077 patients with shoulder instability and rotator cuff injuries, ranging in age from 14 to 85 years. The SST demonstrated acceptable test-retest reliability (intraclass correlation coefficient >0.90) and content validity (floor and ceiling effects <10%). Correlations with the physical functioning component of the Short Form 12 were significant (r = 0.439, P < .05); however, the correlations were not significant when stratified by age group (>60 years) (r = 0.271, P = .349) and injury type (rotator cuff injury) (r = 0.337, P = .085). Correlations with the American Shoulder and Elbow Surgeons were also significant (r = 0.807, P < .001). The construct validity of the SST was acceptable, with all 8 hypotheses demonstrating significance (P < .05). The SST was responsive to change (effect size, 0.81; standardized response mean, 0.81). However, there were differences after stratification for age group and injury type. The SST demonstrated overall acceptable psychometric performance; however, differences were found when data were stratified by age and injury type.

Adolescent↗

Are road safety evaluation studies published in peer reviewed journals more valid than similar studies not published in peer reviewed journals?

The peer review system of scientific journals is commonly assumed to prevent seriously flawed research from getting published. This paper compares the quality of 44 road safety evaluation studies published in peer reviewed journals to the quality of 79 evaluation studies dealing with the same safety measures, but not published in peer reviewed journals, in terms of seven criteria of study validity. Studies were scored for validity in terms of (1) sampling technique, (2) total sample size, (3) mean sample size for each result, (4) specification of accident or injury severity, (5) study design, (6) number of confounding factors controlled and (7) number of moderator variables specified. Confounding factors are all factors that disturb the attribution of a causal relationship between the safety measure being evaluated and the observed changes in safety, moderator variables are all variables that influence the size of the effect of the safety measure. Very few statistically reliable differences in study validity were found between studies published in peer reviewed journals and studies not published in such journals. There was, at best, a weak tendency for studies published in peer reviewed journals to score higher for validity. An interaction was found between author affiliation and type of publication with respect to study validity. Studies published in peer reviewed journals by authors who were at a university scored highest for validity. For a number of reasons, this study must be regarded as exploratory and its results as indicative only. The study does, however, point to a line of research that might be worth pursuing in larger and more rigorous studies.

Automobile Driving↗