Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

The Frail Elderly Functional Assessment questionnaire: its responsiveness and validity in alternative settings.

OBJECTIVE: To test the Frail Elderly Functional Assessment (FEFA) questionnaire for responsiveness (sensitivity to change) to low-level functional tasks in a frail elderly cohort and to evaluate its validity over the telephone or when administered to a caregiver proxy. SUBJECTS: Fifty-eight elderly patients from three urban inpatient rehabilitation settings and an outpatient geriatrics center. METHODS: A prospective, clinical, comparative trial. The FEFA questionnaire was administered serially. For validity, subjects were observed performing the tasks on the questionnaire within 24 hours of each interview. For responsiveness, repeat measures were performed within a 1- to 2-week period. Validity and sensitivity to change (responsiveness) of the questionnaire were determined by correlating patient responses to direct observations by rehabilitation staff. Responsiveness was also determined based on the Guyatt technique that divides clinically significant change by the normal variance, sigma/(2x [mean squared error])1/2, as well as by measures of effect size, standardized response means, and relative efficiency tests for responsiveness. To evaluate FEFA validity in alternative settings, kappa statistic and regression analyses were used based on the previously validated interviewer-administered format. RESULTS: Responsiveness was excellent with effect size (.35), standardized response means (.48), and relative efficiency (2.67) tests as well as Guyatt (1.26). There was 83% agreement when compared with FEFA task performance. Regression between change in FEFA score versus performance testing was significant (r2 = .33; p = .01). ANOVA was significant at a p = .03 for FEFA scores at first measure in rehabilitation compared to second. Correlation for caregiver proxy administration was .92 (p< or =.0001) and for telephone administration was .99 (p<.0001). CONCLUSIONS: The FEFA questionnaire, previously demonstrated to be reliable and valid, is sensitive to functional change (responsive) in frail elderly people. It is also valid when administered by phone or to a caregiver proxy.

Activities of Daily Living↗

Diagnostic validity as a theoretical concept and as a measurable quantity.

The analytical result of a laboratory examination is a scientific fact and has no medical meaning as such. It must be interpreted to become a medical finding. To explain the very complex cognitive procedure of the interpretation a three-level model is used. In an environment of cost containment in health care systems the quality of medical laboratory findings is very important. Analytical results are monitored by quality control procedures. For measuring the performance of medical findings the concept of the 'validity' of a laboratory test is used. Validity means the 'degree of achieving the objective'. Accordingly, a valid laboratory finding is one which correctly answers the question which the physician at the sick-bed directs to the laboratory. Quantitative measures for the validity of interpretation can be developed by an analysis of the underlying classification processes. Characteristic indices describing the validity quantitatively in terms of conditional probabilities can be derived from decision tables. Examples of 'validity indices' are diagnostic (or prognostic) 'sensitivity' and 'specificity'. These indices are powerful tools for developing strategies for the clinical use of laboratory examinations in diagnosis, prognosis and therapy management. Moreover, validity indices are appropriate output quantities for the estimation of effectiveness and efficiency of a diagnostic or prognostic examination.

Clinical Laboratory Techniques↗

The validity of atypical depression in DSM-IV.

Atypical depression has been included in the DSM-IV as an episode specifier of major depressive episodes and dysthymia. This report will review evidence for the clinical validity of atypical depression using operational criteria for the validation of clinical syndromes. English language articles between 1969 and March 1996 were found using a computerized and manual reference search and were selected according to the following criteria: (1) primary research, (2) definition of atypical depression, which includes depression and not anxiety alone, and (3) relevance of data for validation of atypical depression. Studies were evaluated on Kendall's six criteria for establishing clinical validity. There are supporting data for diagnostic validity of atypical depression in the criteria of clinical description and differential treatment response, with atypical depression having a superior response to monoamine oxidase (MAO) inhibitors compared to tricyclic antidepressants. There is still only limited support for the validity of atypical depression in the criteria of pathophysiology, points of rarity with other similar diagnoses, distinctive course and outcome, and genetics. Based on the current evidence, atypical depression is a useful diagnostic concept, particularly for predicting differential drug response, but further research is required to conclusively demonstrate its validity as a clinical syndrome.

Antidepressive Agents↗

Assessing the reliability and validity of the Chinese version of the California Critical Thinking Disposition Inventory.

The purpose of this study was to: (1) translate the California Critical Thinking Disposition Inventory (CCTDI) from English to Chinese; (2) ascertain the reliability and validity of Chinese CCTDI; and (3) assess the psychometric equivalencies across Chinese and English versions of the CCTDI. The CCTDI was designed to measure critical thinking dispositions of truth-seeking, open-mindedness, analyticity, systematicity, inquisitiveness, self-confidence and maturity, which has been approved with a significant difference from prior conceptualizations of critical thinking dispositions. It is the only measurement that has been validated to measure critical thinking disposition and is appropriate for use in nursing. Based on translation theory, the comparative study goal precisely matched with the strategy of decentered translation. The CCTDI was translated in multiple stages and back translated by a panel of bilingual experts. Content validity index (CVI) ranged from 0.50 to 0.80, with an overall CVI of 0.85. Pearson r ranged from 0.33 to 0.79, with an overall correlation of 0.79, indicating that evidence for stability in truth-seeking, open-mindedness and self-confidence existed. To ascertain internal consistency reliability and construct validity, monolingual samples were obtained from 214 and 196 undergraduate nursing students from Taiwan and the USA, respectively. For the Chinese CCTDI, subscale alphas ranged from 0.34 to 0.73, with an overall alpha of 0.71. For the English CCTDI, subscale alphas ranging from 0.52 to 0.73 and an overall alpha of 0.71 were obtained. In terms of a confirmatory factor analysis with LInear Structural RELationships (LISREL), the results indicate that evidence for construct validity existed for truth-seeking, open-mindedness, systematicity, and maturity for the Chinese CCTDI. After allowing some error to exist and deleting three items, evidence for construct validity existed for the remaining subscales. The results of the psychometric equivalencies across Chinese and English CCTDI showed similarity for content validity and reliability for inquisitiveness. In terms of multisample analysis, there were equal forms across all subscales of the two versions. Consequently, although the translation adequacy of the Chinese CCTDI needs to be improved, there is evidence that it is useful for evaluating critical thinking dispositions.

California↗

[Investigation methods in clinical cardiology. IV. Clinical measurements in cardiology: validity and errors of measurements].

Measurements represent an essential part of clinical activity. Very often, however, relevant disagreement in clinical measurements becomes apparent. The sources of this variability are the subjects (patients) that are measured, the measurement instrument itself, and the observer. The assessment of the quality of measurement usually relies on the evaluation of its reproducibility and its validity. The reproducibility is basically measured as the inter-observer concordance, the intra-observer concordance, and the test-retest concordance. The specific parameter used to its quantification (intra-class correlation coefficient, kappa index, graphic methods, etc.) depend on the kind of variable to be measured. The validity of the measurement is the degree to which the measurement is really measuring what we think it should. If an acceptable standard is available, then so called criterion validity is usually assessed. Otherwise the validity should be assessed by other ways that use subjective criteria (content validity and face validity) or empirical criteria (construct validity).

Cardiology↗

Validity of the longitudinal, expert, all data procedure for psychiatric diagnosis in patients with psychoactive substance use disorders.

The longitudinal, expert, all data (LEAD) procedure has been employed as a criterion for the assessment of the procedural validity of diagnostic instruments. This study evaluated the procedure's concurrent, discriminant and predictive validity. Interview and questionnaire data obtained from 100 individuals in a substance abuse treatment program were used to assess current and lifetime substance use disorders and common comorbid disorders. An experienced, doctoral-level clinician formulated LEAD diagnoses for each patient, based on an initial interview, ongoing clinical contact and the results of the research assessment and all available clinical records. LEAD-derived substance use diagnoses showed good concurrent, discriminant and predictive validity. The validity of comorbid diagnoses obtained using the LEAD procedure was generally fair to good. Comparison with diagnoses based only on the clinician's unstructured initial interview showed that the availability of additional data enhanced diagnostic validity. Diagnoses derived by a research technician using the Structured Clinical Interview for DSM-III-R showed validity comparable to that of LEAD diagnoses. To enhance its diagnostic validity, applications of the LEAD standard should include a structured interview. Other variations in the application of the LEAD standard, including a longer evaluation period, may also enhance its performance as a diagnostic criterion measure.

Adult↗

Convergent discriminitive, and predictive validity of the Prostate Cancer Specific Quality of Life Instrument (PROSQOLI) assessment and comparison with analogous scales from the EORTC QLQ-C30 and a trial-specific module. European Organisation for Research and Treatment of Cancer. Core Quality of Life Questionnaire.

The Prostate Cancer Specific Quality of Life Instrument (PROSQOLI) is a measure of health-related quality of life (HRQL) that was designed to be an outcome measure for clinical trials in advanced hormone-resistant prostate cancer. The cross-sectional validity of the PROSQOLI was assessed using baseline data from a randomized trial in which HRQL was also assessed with the European Organisation for Research and Treatment of Cancer (EORTC) Core Quality of Life Questionnaire (QLQ-C30) and a trial-specific quality of life module (QLM-P14). Convergent validity was assessed with the multitrait-multimethod matrix approach; discriminative validity was assessed according to conventional clinical criteria; and predictive validity was assessed by the ability to predict survival duration. These assessments provided strong support for the validity of all PROSQOLI scales except those for family/marriage relationships and passing urine; modifications of these two scales are under evaluation. The strength, consistency, and independence of the prognostic information provided by the HRQL scales were striking. Differences between the instruments were generally subtle. These data support validity of the PROSQOLI and the analogous scales from the QLQ-C30 and QLM-P14 in symptomatic men with advanced hormone resistant prostate cancer. The PROSQOLI is a short, simple, and valid measure of HRQL in this setting.

Analysis of Variance↗

A cross-validation of the European Organization for Research and Treatment of Cancer QLQ-C30 (EORTC QLQ-C30) for Japanese with lung cancer.

The EORTC QLQ-C30 was developed in English-speaking cultures. To determine if this instrument could cross a broad cultural divide and be used in Japan, the cross-cultural validity of its Japanese version was estimated. In evaluating psychometric testing, internal consistency by Cronbach's alpha, item-discrimination by multitrait scaling analysis, and validity analysis with ECOG performance score (PS) and Karnofsky Performance Status Scale (KPS) were performed. The QLQ-C30 (version 1.0) was given to 105 patients with lung cancer. Although the response rate was low in patients with PS 4, the questionnaire was well accepted by patients with PS 0-3. The Japanese QLQ-C30 has a weak scale of role functioning in terms of item discriminative validity. It also has a weak scale of cognitive functioning in items of discriminative validity and internal consistency. However, known-groups comparison showed the expected clinical validity with PS for all the scales except for financial impact, and longitudinally clinical validity with KPS was shown in scales of cognitive functioning, fatigue, and nausea and vomiting. Multitrait scaling analysis showed that the predicted scales constituting quality of life (QOL) in the English-speaking culture were extracted from the Japanese QLQ-C30, and found to be valid in Japan, indicating its possible usefulness as an instrument that is universally applicable across cultures.

Cross-Cultural Comparison↗

Validation of the Insomnia Severity Index as an outcome measure for insomnia research.

Background: Insomnia is a prevalent health complaint that is often difficult to evaluate reliably. There is an important need for brief and valid assessment tools to assist practitioners in the clinical evaluation of insomnia complaints.Objective: This paper reports on the clinical validation of the Insomnia Severity Index (ISI) as a brief screening measure of insomnia and as an outcome measure in treatment research. The psychometric properties (internal consistency, concurrent validity, factor structure) of the ISI were evaluated in two samples of insomnia patients.Methods: The first study examined the internal consistency and concurrent validity of the ISI in 145 patients evaluated for insomnia at a sleep disorders clinic. Data from the ISI were compared to those of a sleep diary measure. In the second study, the concurrent validity of the ISI was evaluated in a sample of 78 older patients who participated in a randomized-controlled trial of behavioral and pharmacological therapies for insomnia. Change scores on the ISI over time were compared with those obtained from sleep diaries and polysomnography. Comparisons were also made between ISI scores obtained from patients, significant others, and clinicians.Results: The results of Study 1 showed that the ISI has adequate internal consistency and is a reliable self-report measure to evaluate perceived sleep difficulties. The results from Study 2 also indicated that the ISI is a valid and sensitive measure to detect changes in perceived sleep difficulties with treatment. In addition, there is a close convergence between scores obtained from the ISI patient's version and those from the clinician's and significant other's versions.Conclusions: The present findings indicate that the ISI is a reliable and valid instrument to quantify perceived insomnia severity. The ISI is likely to be a clinically useful tool as a screening device or as an outcome measure in insomnia treatment research.

Journal Article↗

Reporting of instrument validity and reliability in selected clinical nursing journals, 1989.

Before research findings are applied to practice, the quality of the research must be assessed so that flawed research does not lead inadvertently to flawed practice. Two critical indicators of research quality are the validity and reliability of the data collection instruments. This article summarizes the principles of instrument validity and reliability and identifies deviations from these principles in a random sample of 55 research studies published in 1989 in five refereed nursing journals targeted toward practicing clinicians. Using a valid and reliable instrument, the investigators found that even with a policy of giving authors "the benefit of the doubt," 47% of the research studies contained no evidence of validity for any data collection instruments and 36% had no evidence of reliability; 29% had no evidence of either validity or reliability. Content validity, a basic requirement for all research instruments, was addressed in only 27% of the studies. This article provides documentation, justification, and suggestions for nursing educators, journal editors, and researchers to take action to improve the reporting of instrument validity and reliability to help ensure the quality of the research on which nursing practice is based.

Clinical Nursing Research↗

Added value for tandem mass spectrometry shotgun proteomics data validation through isoelectric focusing of peptides.

A very popular approach in proteomics is the so-called "shotgun LC-MS/MS" strategy. In its mostly used form, a total protein digest is separated by ion exchange fractionation in the first dimension followed by off- or on-line RP LC-MS/MS. We replaced the first dimension by isoelectric focusing in the liquid phase using the Off-Gel device producing 15 fractions. As peptides are separated by their isoelectric point in the first dimension and hydrophobicity in the second, those experimentally derived parameters (pI and R(T)) can be used for the validation of potentially identified peptides. We applied this strategy to a cellular extract of Drosophila Kc167 cells and identified peptides with two different database search engines, namely PHENYX and SEQUEST, with PeptideProphet validation of the SEQUEST results. PHENYX returned 7582 potential peptide identifications and SEQUEST 7629. The SEQUEST results were reduced to 2006 identifications by validation with PeptideProphet. Validation of the PeptideProphet, SEQUEST and PHENYX results by pI and R(T) parameters confirmed 1837 PeptideProphet identifications while in the remainder of the SEQUEST results another 1130 peptides were found to be likely hits. The validation on PHENYX resulted in the fixation of a solid p-value threshold of <1 x 10(-04) that sets by itself the correct identification confidence to >95%, and a final count of 2034 highly confident peptide identifications was achieved after pI and R(T) validation. Although the PeptideProphet and PHENYX datasets have a very high confidence the overlap of common identifications was only at 79.4%, to be explained by the fact that data interpretation was done searching different protein databases with two search engines of different algorithms. The approach used in this study allowed for an automated and improved data validation process for shotgun proteomics projects producing MS/MS peptide identification results of very high confidence.

Algorithms↗

Using the EuroQoI 5-D in the Catalan general population: feasibility and construct validity.

Spanish and Catalan versions of the EuroQoi 5-D (EQ-5D) were included in the Catalan Health Interview Survey (CHIS) and administered to a randomly selected cross-section of 12,245 individuals from the Catalan general population. This paper analyses the feasibility, convergent validity and construct validity of three parts of the EQ-5D (the descriptive system, the visual analogue scale (VAS) and the Spanish tariff) using the results obtained in the CHIS. The feasibility was assessed by the number of missing responses. The convergent validity was based on the correlations between the EQ-5D scores and the scores on the General Health Questionnaire (GHQ) and on an index of self-perceived overall health. The construct validity was assessed by analysing the degree to which lower scores on the EQ-5D correlated positively with increasing age, being female, being in a lower social class or having a lower level of education and with increasing levels of disability, co-morbidity, restricted activity, mental health problems and poor self-perceived health. A low number of missing responses on the descriptive system and the VAS (1.5%) indicated a high level of acceptance. A marked ceiling effect was found, with 67% of the sample reporting no problem in any EQ dimension. The convergent validity with the GHQ was generally low, though moderate on the mood dimension. Self-perceived overall health correlated moderately to strongly with the mean VAS and tariff values. The positive correlations between lower scores on all three elements of the EQ-5D and increasing age, increasing levels of disability, comorbidity, restricted activity, mental health problems and poor self-perceived health provide some evidence of the instrument's construct validity, as does the fact that women reported more problems than men. Multivariate analyses using the VAS and tariff values as dependent variables and all of the sociodemographic and health variables as independent variables reached R2 values of 0.45 and 0.81, respectively. The Spanish and Catalan versions of the EQ-5D have proved to be feasible and valid for use in health interview surveys.

Adolescent↗

Validation subset selections for extrapolation oriented QSPAR models.

One of the most important features of QSPAR models is their predictive ability. The predictive ability of QSPAR models should be checked by external validation. In this work we examined three different types of external validation set selection methods for their usefulness in in-silico screening. The usefulness of the selection methods was studied in such a way that: 1) We generated thousands of QSPR models and stored them in 'model banks'. 2) We selected a final top model from the model banks based on three different validation set selection methods. 3) We predicted large data sets, which we called 'chemical universe sets', and calculated the corresponding SEPs. The models were generated from small fractions of the available water solubility data during a GA Variable Subset Selection procedure. The external validation sets were constructed by random selections, uniformly distributed selections or by perimeter-oriented selections. We found that the best performing models on the perimeter-oriented external validation sets usually gave the best validation results when the remaining part of the available data was overwhelmingly large, i.e., when the model had to make a lot of extrapolations. We also compared the top final models obtained from external validation set selection methods in three independent and different sizes of 'chemical universe sets'.

Computer Simulation↗

Development and validation of the Eating Disorder Diagnostic Scale: a brief self-report measure of anorexia, bulimia, and binge-eating disorder.

This article describes the development and validation of a brief self-report scale for diagnosing anorexia nervosa, bulimia nervosa, and binge-eating disorder. Study 1 used a panel of eating-disorder experts and provided evidence for the content validity of this scale. Study 2 used data from female participants with and without eating disorders (N = 367) and suggested that the diagnoses from this scale possessed temporal reliability (mean kappa = .80) and criterion validity (with interview diagnoses; mean kappa = .83). In support of convergent validity, individuals with eating disorders identified by this scale showed elevations on validated measures of eating disturbances. The overall symptom composite also showed test-retest reliability (r = .87), internal consistency (mean alpha = .89), and convergent validity with extant eating-pathology scales. Results implied that this scale was reliable and valid in this investigation and that it may be useful for clinical and research applications.

Adolescent↗

Introduction to the special section on incremental validity and utility in clinical assessment.

This special section focuses on the incremental validity and utility of clinical assessment data. The lack of replicated incremental validity research limits the ability of psychologists to establish their assessment practices on a solid empirical footing. The articles in this section deal with conceptual, methodological, and content issues in the development and use of clinical measures. The authors addressed several aspects of incremental validity research, including (a) the evaluation of the magnitude of validity increments, (b) the use of incremental validity data for test development and validation, (c) the costs of assessment, (d) the use of multi-informant and multimethod assessments, and (e) the treatment utility of assessment. Obstacles limiting the use of incremental validity research to inform and guide clinical practice are also emphasized.

Humans↗

The incremental validity of psychological testing and assessment: conceptual, methodological, and statistical issues.

There has been insufficient effort in most areas of applied psychology to evaluate incremental validity. To further this kind of validity research, the authors examined applicable research designs, including those to assess the incremental validity of test instruments, of test-informed clinical inferences, and of newly developed measures. The authors also considered key statistical and measurement issues that can influence incremental validity findings, including the entry order of predictor variables, how to interpret the size of a validity increment, and possible artifactual effects in the criteria selected for incremental validity research. The authors concluded by suggesting steps for building a cumulative research base concerning incremental validity and by describing challenges associated with applying nomothetic research findings to individual clinical cases.

Humans↗

Cross-validatory selection of test and validation sets in multivariate calibration and neural networks as applied to spectroscopy.

Cross-validated and non-cross-validated regression models using principal component regression (PCR), partial least squares (PLS) and artificial neural networks (ANN) have been used to relate the concentrations of polycyclic aromatic hydrocarbon pollutants to the electronic absorption spectra of coal tar pitch volatiles. The different trends in the cross-validated and non-cross-validated results are discussed as well as a method for the production of a true cross-validated neural network regression model. It is shown that the methods must be compared through the errors produced in the validation sets as well as those given for the final model. Various methods for calculation of errors are described and compared. The separation of training, validation and test sets into fully independent groups is emphasized. PLS outperforms PCR using all indicators. ANNs are inferior to multivariate techniques for individual compounds but are reasonably effective in predicting the sum of PAHs in the mixture set.

Calibration↗

Reliability and validity of the nutritional form for the elderly (NUFFE).

AIM: The aim of this study was to test the reliability and validity of the Nutritional Form for the Elderly (NUFFE). BACKGROUND: The prevalence of undernutrition among older people in nursing homes and hospitals reaches high levels. Assessment of older patients' nutritional status is an important task for nurses in clinical care. To use a simple nutritional assessment instrument for older people is one approach for nurses. Examples of such instruments are the well validated Mini Nutritional Assessment (MNA) and the newly developed NUFFE. METHODS: A total of 114 consecutively chosen, newly admitted older patients in an elder care rehabilitation ward in western Sweden were interviewed using the NUFFE and MNA. Arm and calf circumferences, body mass index (BMI), and presence of pressure sores and skin ulcers were noted as part of the MNA on admission. Weight was monitored and BMI calculated on discharge. Serum albumin levels on admission and discharge were used if these were available in the records. Reliability of the NUFFE was measured as homogeneity. Criterion related validity, concurrent validity, construct validity, and predictive validity were assessed with different statistical methods. The regional research ethics committee approved the study. RESULTS: The results showed that the NUFFE is a fairly reliable and valid instrument for identifying actual and potential undernutrition among older patients. CONCLUSION: The NUFFE is a simple tool for nurses to use to assess older patients with the aim of detecting undernourished individuals and those at risk for undernutrition. When doing a nutritional assessment with the NUFFE, the BMI ought also to be calculated. The assessment could also be combined with food intake recording for a period of time.

Aged↗