Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A methodology for validating numerical ground water models.

Ground water validation is one of the most challenging issues facing modelers and hydrogeologists. Increased complexity in ground water models has created a gap between model predictions and the ability to validate or build confidence in predictions. Specific procedures and tests that can be easily adapted and applied to determine the validity of site-specific ground water models do not exist. This is true for both deterministic and stochastic models, with stochastic models posing the more difficult validation problem. The objective of this paper is to propose a general validation approach that addresses important issues recognized in previous validation studies, conferences, and symposia. The proposed method links the processes for building, calibrating, evaluating, and validating models in an iterative loop. The approach focuses on using collected validation data to reduce uncertainty in the model and narrow the range of possible outcomes. This method is designed for stochastic numerical models utilizing Monte Carlo simulation approaches, but it can be easily adapted for deterministic models. The proposed methodology relies on the premise that absolute validity is not theoretically possible, nor is it a regulatory requirement. Rather, the proposed methodology highlights the importance of testing various aspects of the model and using diverse statistical tools for rigorous checking and confidence building in the model and its predictions. It is this confidence that will encourage regulators and the public to accept decisions based on the model predictions. This validation approach will be applied to a model, described in this paper, dealing with an underground nuclear test site in rural Nevada.

Environmental Monitoring↗

Validity of a self-administered food frequency questionnaire (FFQ) and its generalizability to the estimation of dietary folate intake in Japan.

BACKGROUND: In an epidemiological study, it is essential to test the validity of the food frequency questionnaire (FFQ) for its ability to estimate dietary intake. The objectives of our study were to 1) validate a FFQ for estimating folate intake, and to identify the foods that contribute to inter-individual variation of folate intake in the Japanese population. METHODS: Validity of the FFQ was evaluated using 28-day weighed dietary records (DRs) as gold standard in the two groups independently. In the group for which the FFQ was developed, validity was evaluated by Spearman's correlation coefficients (CCs), and linear regression analysis was used to identify foods with large inter-individual variation. The cumulative mean intake of these foods was compared with total intake estimated by the DR. The external validity of the FFQ and intake from foods on the same list were evaluated in the other group to verify generalizability. Subjects were a subsample from the Japan Public Health Center-based prospective Study who volunteered to participate in the FFQ validation study. RESULTS: CCs for the internal validity of the FFQ were 0.49 for men and 0.29 and women, while CCs for external validity were 0.33 for men and 0.42 for women. CCs for cumulative folate intake from 33 foods selected by regression analysis were also applicable to an external population. CONCLUSION: Our FFQ was valid for and generalizable to the estimation of folate intake. Foods identified as predictors of inter-individual variation in folate intake were also generalizable in Japanese populations. The FFQ with 138 foods was valid for the estimation of folate intake, while that with 33 foods might be useful for estimating inter-individual variation and ranking of individual folate intake.

Diet↗

Evolution of a companywide validation program.

A successful computer validation program requires a solid foundation of policies, guidelines, and procedures. An unstructured approach to validation may yield adequate results for a specific computerized system, but it will not sustain ongoing, consistent computer validation activities over time. The evolution of a strong computer validation program starts with an awareness of the need for such a program and builds on that established base. One approach is a step-by-step method in which each new step builds on the results of one or more of the previous steps. However, the order of events is not as important as the ultimate completion of all the building blocks. A strong computer validation program requires a companywide policy and an administrator who provides oversight and chairs the computer validation committee. The committee should be established to provide departmental leadership and to develop company guidelines for computer validation. Committee members will also create system inventories and classify and prioritize those systems in terms of the need for validation. Departmental standard operating procedures must be developed to provide standardized methods for routine validation activities. Finally the quality assurance unit should have procedures in place that provide for its involvement in the computer validation process.

Computer Systems↗

Validation of an outcomes instrument for tonsil and adenoid disease.

OBJECTIVE: To design and validate a disease-specific health status instrument-the Tonsil and Adenoid Health Status Instrument-for use in children with tonsil and adenoid disease. DESIGN: Prospective psychometric and clinimetric instrument validation in 3 stages. SETTINGS: A tertiary academic pediatric specialty hospital and a tertiary academic hospital, in 2 different cities. PATIENTS/OTHER PARTICIPANTS: Children with tonsil and adenoid disease presenting for evaluation and treatment (n = 224). INTERVENTION/METHOD: Prospective instrument validation. Stage 1 consisted of initial item testing, reduction, and subscale construction; stage 2, reliability and validity testing, factor analysis, and final item reduction; and stage 3, responsiveness analysis. MAIN OUTCOME MEASURES: Test-retest and internal consistency reliability; content, construct, and criterion validity; orthogonal principal components factor analysis; and response sensitivity analysis. RESULTS: Factor analysis and item analysis confirmed 6 distinct subscales measuring different constructs (aspects) of disease-specific health status that are affected by tonsil and adenoid disease: eating and swallowing, airway and breathing, infections, health care utilization, cost of care, and behavior. For each subscale, the Tonsil and Adenoid Health Status Instrument demonstrated excellent test-retest reliability (r = 0.72-0.88) and internal consistency reliability (Cronbach alpha = .73-.87). Content validity was ensured during the design process. Construct validity was demonstrated by means of convergent and divergent validity with a global quality-of-life instrument (the Child Health Questionnaire, version PF28). Criterion validity was also satisfactory. Finally, the instrument was appropriately sensitive, with high standardized response means and effect sizes. CONCLUSIONS: The Tonsil and Adenoid Health Status Instrument is a valid, reliable, and sensitive instrument with 6 distinct subscales. This instrument has significant utility for outcomes research in children with tonsil and adenoid disease.

Adenoids↗

The importance of validating the diagnosis of coronary heart disease when measuring secondary prevention: a cross-sectional study in general practice.

PURPOSE: To compare levels of recorded risk factors and drug treatment between patients with validated and non-validated diagnoses of coronary heart disease (CHD) in Northern Ireland. METHODS: Patients with a nitrate prescription in the previous year or a CHD Read code were identified from computer records of 25 practices, stratified by partnership size and area board. Computer and paper records of a random sample of 10% of these were searched for specified criteria to validate the diagnosis of CHD. The diagnosis was considered valid if the patient was found to have one or more positive investigations for CHD. Records of blood pressure, cholesterol, blood sugar, body mass index and drugs prescribed were taken into account. RESULTS: The combined practice population was 151,071; 7338 (4.86%) were identified by the computer search as meeting the defined entry criteria for CHD. Among the 10% random sample the diagnosis of CHD could not be validated for 36.5% (265/727). Significantly more patients with a validated than non-validated diagnosis had recorded cholesterol levels below 5.0 mmol/l (55.8 vs. 34.5%, p < 0.001) and were prescribed aspirin (75.3 vs. 40.8%, p < 0.001), beta-blockers (51.5 vs. 28.3%, p < 0.001), angiotensin-converting-enzyme inhibitors (29.2 vs. 15.5%, p < 0.001) and lipid-lowering drugs (50.9 vs. 23.0%, p < 0.001). A recent nitrate prescription had a higher predictive value for validated CHD than a Read code for CHD alone (71.2 vs. 53.1%, p < 0.001). No other significant differences were found between the two groups regarding the extent or levels of recorded risk factors. CONCLUSIONS: Patients with a validated diagnosis of CHD appear to be better managed than those whose diagnosis has not been confirmed. Validation of diagnosis has important implications for assessing the provision of secondary prevention and for clinical governance.

Adolescent↗

Constructing and Validating Motive Bridging Inferences

Understanding Jane left early for the birthday party, She spent an hour shopping at the mall requires detecting that the first statement motivates the second. The validation model states that before accepting this bridging inference, the reader validates it with reference to relevant knowledge. In particular, a mediating idea is first derived from the text outcome and its candidate motive. If the mediating idea is supported by general knowledge, then the inference has been validated. In tests of this anaylsis, experimental subjects read motive or control sequences and then answered questions probing the knowledge hypothesized to validate the motive inferences, such as Do birthday parties involve presents? Five experiments confirmed that understanding motive sequences facilitates validating knowledge. A control procedure also refuted a priming counterexplanation of these effects (Experiment 1). Validation processing obtained for motive-outcome statements separated by two to four sentences in coherent sequences (Experiments 2 to 4). Inferred and explicit validating knowledge had a similar representational status (Experiment 3). Whereas proofreading abolished the validation effect, a reading strategy promoting causal processing did not enhance it (Experiment 4). A delayed priming procedure indicated that validating knowledge is integrated with the text representation (Experiment 5). The implications of these findings for the constructionist and minimal inference analyses were explored. The validation effects were simulated using construction-integration model.

Journal Article↗

Validation of a new basic virtual reality simulator for training of basic endoscopic skills: the SIMENDO.

BACKGROUND: The aim of this study was to establish content, face, concurrent, and the first step of construct validity of a new simulator, the SIMENDO, in order to determine its usefulness for training basic endoscopic skills. METHODS: The validation started with an explanation of the goals, content, and features of the simulator (content validity). Then, participants from eight different medical centers consisting of experts (> or =100 laparoscopic procedures performed) and surgical trainees (<100) were informed of the goals and received a "hands-on tour" of the virtual reality (VR) trainer. Subsequently, they were asked to answer 28 structured questions about the simulator (face validity). Ratings were scored on a scale from 1 (very bad/useless) to 5 (excellent/very useful). Additional comments could be given as well. Furthermore, two experiments were conducted. In experiment 1, aimed at establishing concurrent validity, the training effect of a single-handed hand-eye coordination task in the simulator was compared with a similar task in a conventional box trainer and with the performance of a control group that received no training. In experiment 2 (first step of construct validity), the total score of task time, collisions, and path length of three consecutive runs in the simulator was compared between experts (>100 endoscopic procedures) and novices (no experience). RESULTS: A total of 75 participants (36 expert surgeons and 39 surgical trainees) filled out the questionnaire. Usefulness of tasks, features, and movement realism were scored between a mean value of 3.3 for depth perception and 4.3 for appreciation of training with the instrument. There were no significant differences between the mean values of the scores given by the experts and surgical trainees. In response to statements, 81% considered this VR trainer generally useful for training endoscopic techniques to residents, and 83% agreed that the simulator was useful to train hand-eye coordination. In experiment 1, the training effect for the single-handed task showed no significant difference between the conventional trainer and the VR simulator (concurrent validity). In experiment 2, experts scored significantly better than novices on all parameters used (construct validity). CONCLUSION: Content, face, and concurrent validity of the SIMENDO is established. The simulator is considered useful for training eye-hand coordination for endoscopic surgery. The evaluated task could discriminate between the skills of experienced surgeons and novices, giving the first indication of construct validity.

Adult↗

Approaches for assessing the validity of a functional observational battery.

As neurobehavioral assessments during the preliminary stages of chemical testing are more widely undertaken, it is critical that the screening procedures utilized be valid indicators of neurobehavioral function and that they be sensitive, specific, and reliable. Efforts in this laboratory have been directed towards assessing these features in the use of a functional observational battery (FOB). For the purpose of assessing validity, we have examined FOB data which addresses the issues of criterion, predictive, concurrent, and construct validities. The FOB appears to be valid for detecting chemical-induced neurological dysfunction in rats, i.e., shows a good degree of criterion validity. Furthermore, in many instances the effects observed with the FOB may be predictive of symptomatology in humans. When comparisons can be made between effects detected with the FOB and other methods of measuring neurotoxicity (e.g., neuropathology), concurrent validity can also be established. To assess construct validity, effects of neurotoxicants can be classified into functional domains which are described by various measures in the FOB. Approaches for assessing the validity of the test method thus include answering specific research questions directed at assessing criterion, predictive, concurrent, and construct validity. Available data indicate that, in these aspects, the FOB is a valid screening method for the detection of neurotoxicity.

Animals↗

The validation of three human reliability quantification techniques--THERP, HEART and JHEDI: Part III--Practical aspects of the usage of the techniques.

This is the third paper in a series of three dealing with the detailed investigation of the empirical validity of three human reliability assessment (HRA) techniques. The first paper introduced the need for validation and specified the three techniques most requiring validation. The second paper detailed the results of an extensive independent validation experiment. This experimental validation involved 30 UK assessors using the techniques THERP, HEART and JHEDI (10 assessors per technique) to estimate the human error probabilities (HEPs) for 30 nuclear power and reprocessing (NP&R) tasks. The results for all three techniques were positive in terms of significant correlations, and general precision levels of 72% of all HEP estimates within a factor of 10 of the true value (unknown to the assessors). These results lend support to the empirical validity of these techniques in particular, and to HRA in general. However, the results were not all positive. In particular the consistency of usage of the techniques was variable. Additionally, subjects were generally not good at knowing their own uncertainty, i.e. they were not able to accurately predict when they were accurate nor when they were inaccurate. This desirable parameter is known as calibration, and the results from the validation suggested that subjects were not well-calibrated. This paper aims to determine how consistency of usage can be improved and to discern whether certain task types are, in practice, not well-assessed by the techniques, and hence are effectively currently beyond these techniques' abilities. Such information is aimed at aiding the HRA practitioner, or the ergonomist, interested in using these techniques. Recommendations for improving calibration are also discussed in this paper. A subsidiary but important focus of this paper is of a more fundamental nature, and of more general interest to the ergonomist. It concerns the validity of the techniques from an error reduction perspective. Currently these techniques may be used to identify how to reduce error probability, which is generally (in the qualitative sense) within the domain of ergonomics. One major mechanism for HRA-based error reduction is the utilisation of Performance Shaping Factor (PSF) information. This paper considers the validity of these PSF as ergonomics constructs. Drawing results from the validation exercise, it is seen how different PSF can be applied to the same scenario and can result in the same error probability, but will result in different error reduction guidance. It is therefore recommended that error reduction guidance must be based on a composite analysis of the results of the task, error identification and quantification analyses, with most weighting given to the qualitative analyses.

Evaluation Studies as Topic↗

Teasing apart quality and validity in systematic reviews: an example from acupuncture trials in chronic neck and back pain.

The objectives of the study were (1) to carry out a systematic review to assess the analgesic efficacy and the adverse effects of acupuncture compared with placebo for back and neck pain and (2) to develop a new tool, the Oxford Pain Validity Scale (OPVS), to measure validity of findings from randomized controlled trials (RCTs), and to enable ranking of trial findings according to validity within qualitative reviews. Published RCTs (of acupuncture at both traditional and non-traditional points) were identified from systematic searching of bibliographic databases (e.g. MEDLINE) and reference lists of retrieved reports. Pain outcome data were extracted with preference given to standardized outcomes such as pain intensity. Information on adverse effects was also extracted. All included trials were scored using a five-item 0-16 point validity scale (OPVS). The individual RCTs were ranked according to their OPVS score to enable more weight to be placed on the trials of greater validity when drawing an overall conclusion about the efficacy of acupuncture for relieving neck and back pain. Statistical analyses were carried out on the OPVS scores to assess the relationship between trial finding (positive or negative) and validity. Thirteen RCTs met the inclusion criteria. Five trials concluded that acupuncture was effective, and eight concluded that it was not effective for relieving back or neck pain. There was no obvious difference between the findings of trials using traditional and non-traditional points. Using the new OPVS scale, the validity scores of the included trials ranged from 4 to 14. There was no significant relationship between OPVS score and trial finding (positive versus negative). Authors' conclusions did not always agree with their data. We drew our own conclusions (positive/negative) based on the data presented in the reports. Re-analysis using our conclusions showed a significant relationship between OPVS score and trial finding, with higher validity scores associated with negative findings. OPVS is a useful tool for assessing the validity of trials in qualitative reviews. With acupuncture for chronic back and neck pain, we found that the most valid trials tended to be negative. There is no convincing evidence for the analgesic efficacy of acupuncture for back or neck pain.

Acupuncture Therapy↗

Plateletapheresis: instrumentation validation.

Plateletapheresis instrumentation validation is required to document that a new or modified instrument or technique is capable of consistently producing acceptable products at the production center using their equipment, personnel, and counting techniques even though the instrument or technique may already have FDA or equivalent approval for use. To pursue the process of validation, several questions need to be addressed: when is it required, what products are validated, what parameters are monitored, and how many products are required. Validation is required when a new instrument or technique (process) is used that could affect the quality of the product. According to the FDA, each apheresis system (e.g., Spectra LRS, Amicus) and each type of product (e.g., single, double, triple) need to be validated separately. Parameters to be validated vary, but usually platelet (plt) yield, white blood cell (WBC) content (if products are labeled "leukoreduced"), and 5-day storage pH are monitored. The number of procedures monitored is also quite variable, but we use 20 samples for highly variable parameters such as platelet yield and WBC content and five samples for less variable parameters such as 5-day storage pH. As an example, we validated the Fenwal Amicus (Baxter Biotech) for single apheresis platelet products. With 20 samples, we found that: 85% of the products contained > or = 3 x 10(11) plt (requirement was at least 75% contain > or = 3 x 10(11) plt); platelet concentration of all products was < or = 1.515 x 10(6) plt/microL (requirement was < or = 2.435 x 10(6) plt/microL), and WBC content was < 1 x 10(6) WBC in all products (requirement was all products contain < 5 x 10(6) WBC). In addition, in five samples, the 5-day storage pH was 6.89-7.25 (requirement was all products should be > or = 6.2 pH). Once validation is complete and acceptable, the process should be monitored on a regular basis using some form of process control. Statistical process control programs are available that can assist in documenting validation and ongoing process control. With the use of process validation and ongoing process control, the plateletapheresis center can assure that acceptable products are consistently being produced.

Forms and Records Control↗

Incremental validity of new clinical assessment measures.

The authors address conceptual and methodological foundations of incremental validity in the evaluation of newly developed clinical assessment measures. Incremental validity is defined as the degree to which a measure explains or predicts a phenomenon of interest, relative to other measures. Incremental validity can be evaluated on several dimensions, such as sensitivity to change, diagnostic efficacy, content validity, treatment design and outcome, and convergent validity. Indices of incremental validity can vary depending on the criterion measures, comparison measures, and individual differences in samples. The authors review the rationale for, principles, and methods of incremental validation, including the selection of comparison and criterion measures, and address data analytic strategies and the conditional nature of incremental validity evaluations in the selection of measures. Incremental validity contributes to, but is different from, cost-benefits, which reflect the cost of acquiring the data and the benefits from the data. The impact of an incremental validity index on whether a measure is selected will be moderated by the cost of acquiring the new data, the importance of the measured phenomenon, and the clinical utility of the new data.

Humans↗

Experience with two validation methods in a prevalence survey on nosocomial infections.

OBJECTIVE: To determine whether an investigator effect remained on the first German study on the prevalence of nosocomial infections Nosokomiale Infektionen in Deutschland Erfassung und Prävention (NIDEP), despite extensive validation efforts. DESIGN: Two validation methods were applied: bedside validation and validation by case studies. In both cases, the results of the four investigators were compared with the diagnosis of gold standard observers. SETTING: Validation measures were applied before, intermittently, during, and at the end of the surveillance period in 72 acute-care hospitals with 14,966 patients. RESULTS: The overall sensitivity in the bedside-validation periods was 89.0%; the overall specificity was 99.5%. For validation by case studies, overall sensitivity was 95.6%, and overall specificity was 92.8%. At the end of the surveillance, a remarkable investigator effect was found. CONCLUSION: Despite validation results that were assessed as satisfactory, based on available literature, an investigator effect was observed. This underlines the need for data validation and the formulation of recommendations for data validation. Clarification of the Centers for Disease Control and Prevention criteria for pneumonia and primary bloodstream infection and the inclusion of some diagnostic test results may reduce or prevent an investigator effect in future studies.

Bias↗

How valid and reliable are patient satisfaction data? An analysis of 195 studies.

OBJECTIVE: To assess the properties of validity and reliability of instruments used to assess satisfaction in a broad sample of health service user satisfaction studies, and to assess the level of awareness of these issues among study authors. DESIGN: Examination and analysis of 195 papers published in 1994 in 139 journals. The following databases were searched: British Nursing Index, CINAHL, EMBASE, MedLine, Popline, and PsycLIT. MAIN MEASURES: Number and types of strategies used for content, criterion, and construct validity, and for stability and internal consistency. Associations between validity/reliability and other study characteristics. RESULTS: Eighty-nine (46%) of the 195 studies reported some validity or reliability data; 76 reported some element of content validity; 14 reported criterion validity, with patient's intent to return the most commonly used criterion; four reported construct validity. Thirty-four studies reported internal consistency reliability, 31 of which used Cronbach's coefficient alpha; eight studies reported test-retest reliability. Only 11 studies (6% of the 181 quantitative studies) reported content validity and criterion or construct validity and reliability. 'New' instruments designed specifically for the reported study demonstrated significantly less evidence for reliability/validity than did 'old' instruments. CONCLUSION: With few exceptions, the study instruments in this sample demonstrated little evidence of reliability or validity. Moreover, study authors exhibited a poor understanding of the importance of these properties in the assessment of satisfaction. Researchers must be aware that this is poor research practice, and that lack of a reliable and valid assessment instrument casts doubt on the credibility of satisfaction findings.

Data Collection↗

Testing for predictive validity in health care education research: a critical review.

BACKGROUND: The assessment of predictive validity is the essential core from which a sound model of prediction is built. METHOD: Three methods for assessing predictive validity in health care education research were reviewed: longitudinal profile development, cross-validation, and inspection of the adjusted R2. A total of 47 articles published between 1973 and 1993 in nine health care disciplines were critically reviewed to determine whether the studies tested for predictive validity by using these methods. RESULTS: Very few of the 47 studies used at least one of the three methods for assessing predictive validity. Furthermore, the proportion of variance explained that is reported in the articles is typically small even before assessment of predictive validity. It is sobering to note that these small values may be inflated, since shrinkage is likely to occur when assessing predictive validity on a second, or cross-validation, sample. CONCLUSION: The scarcity of testing of predictive validity in the studies reviewed highlights the necessity of future research to establish the degree of predictive validity, if improvements in predicting success in health care education research are to be realized.

Education↗

The reliability and validity of three dimensional ultrasound volumetric measurements using an in vitro balloon and in vivo uterine model.

OBJECTIVE: To evaluate the reliability and validity of two and three dimensional ultrasound volumetric measurements using balloon and uterine models. DESIGN: Prospetive observational study. SETTING: Obstetric ultrasound department at a university teaching hospital. METHOD: Two and three dimensional ultrasound volumetric measurements (with 5, 10 and 15 ultrasonic slices) were performed on 30 different sets of ultrasound images obtained from 15 water filled balloons with volumes ranging from 19 to 697mL. The measurements were performed independently by two observers who were blinded to the true volumes of the balloons. For the uterine model, only three dimensional ultrasonic volume measurements were performed independently on 16 uteri by two observers who were again unaware of the definitive uterine volumes. OUTCOME MEASURE: For the assessment of intra-and inter-rater reliability, the intraclass correlation coefficient was used. The index of concordance between the ultrasonic volumes and those obtained by the reference standard (validity) was assessed with the conventional Pearson's correlation coefficient, limits of agreement method and the intra-class correlation coefficient. RESULTS: High levels of reliability and validity were obtained for both two and three dimensional ultrasound balloon volume measurements. For two dimensional ultrasonic volume measurements, the intra-class correlation coefficient ranged from 0.992 to 0.998 for reliability and validity whereas the Pearson's correlation coefficient for validity was 0.996. With three dimensional ultrasonic volume measurements, the intra-class correlation coefficient ranged from 0.991 to 0.999 for reliability and validity whereas the Pearson's correlation coefficient for validity was 0.999. Both two and three dimensional ultrasonic measurements tended to underestimate the true balloon volume with the largest observed mean difference obtained with three dimensional ultrasound measurements using five ultrasonic slices and the smallest value obtained with three dimensional ultrasound measurements employing 15 ultrasonic slices. The mean difference in volume measurement for two dimensional ultrasound was intermediate between these two values. However, two dimensional ultrasound volume measurement generated the largest range between the limits of agreement whereas the smallest range was obtained with three dimensional ultrasound using 10 ultrasonic slices. The intra-class correlation coefficient for reliability and validity with three dimensional ultrasonic uterine volume estimation ranged from 0.956 to 0.996 whereas the Pearson's correlation coefficient for validity ranged from 0.993 to 0.999). The use of three dimensional ultrasound also consistently under-estimated the actual uterine volumes. The larger the number of ultrasonic slices employed for three dimensional ultrasound, the smaller was the mean difference between the ultrasonic and true uterine volume measurements and the smaller the limits of agreement. CONCLUSIONS: The reliability and validity of balloon and uterine volume measurement by three dimensional ultrasound is high. This allows further research on three dimensional ultrasound for measuring pelvic organ volumes in the prediction of pelvic pathology.

Female↗

External validity, generalizability, and knowledge utilization.

PURPOSE: To examine the concepts of external validity and generalizability, and explore strategies to strengthen generalizability of research findings, because of increasing demands for knowledge utilization in an evidence-based practice environment. FRAMEWORK: The concepts of external validity and generalizability are examined, considering theoretical aspects of external validity and conflicting demands for internal validity in research designs. Methodological approaches for controlling threats to external validity and strategies to enhance external validity and generalizability of findings are discussed. CONCLUSIONS: Generalizability of findings is not assured even if internal validity of a research study is addressed effectively through design. Strict controls to ensure internal validity can compromise generalizability. Researchers can and should use a variety of strategies to address issues of external validity and enhance generalizability of findings. Enhanced external validity and assessment of generalizability of findings can facilitate more appropriate use of research findings.

Data Collection↗

Validation of the Rockall risk scoring system in upper gastrointestinal bleeding.

BACKGROUND: Several scoring systems have been developed to predict the risk of rebleeding or death in patients with upper gastrointestinal bleeding (UGIB). These risk scoring systems have not been validated in a new patient population outside the clinical context of the original study. AIMS: To assess internal and external validity of a simple risk scoring system recently developed by Rockall and coworkers. METHODS: Calibration and discrimination were assessed as measures of validity of the scoring system. Internal validity was assessed using an independent, but similar patient sample studied by Rockall and coworkers, after developing the scoring system (Rockall's validation sample). External validity was assessed using patients admitted to several hospitals in Amsterdam (Vreeburg's validation sample). Calibration was evaluated by a chi2 goodness of fit test, and discrimination was evaluated by calculating the area under the receiver operating characteristic (ROC) curve. RESULTS: Calibration indicated a poor fit in both validation samples for the prediction of rebleeding (p<0.0001, Vreeburg; p=0.007, Rockall), but a better fit for the prediction of mortality in both validation samples (p=0.2, Vreeburg; p=0.3, Rockall). The areas under the ROC curves were rather low in both validation samples for the prediction of rebleeding (0.61, Vreeburg; 0.70, Rockall), but higher for the prediction of mortality (0.73, Vreeburg; 0.81, Rockall). CONCLUSIONS: The risk scoring system developed by Rockall and coworkers is a clinically useful scoring system for stratifying patients with acute UGIB into high and low risk categories for mortality. For the prediction of rebleeding, however, the performance of this scoring system was unsatisfactory.

Adolescent↗