Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Quality data: what are they?

Nowadays, quality has become a very important factor in almost all areas of endeavour. The data generated from tests for the assessment of potentially toxic chemicals is obviously no exception. It is necessary, therefore, that quality systems be developed to ensure that the data generated to support these tests are of good quality. An acceptable quality system should require that, where applicable, the tests be performed according to defined guidelines. Once defined guidelines have been identified for the type of test to be performed, it is then necessary to design a plan which describes how, when, where and by whom the data will be generated. If at all possible, the data should be generated according to written standard procedures which provide for the production of data to the same quality standard. The data should be generated and collected by properly trained staff using data collection systems (paper or electronic media) which ensure the accuracy, reliability and integrity of the data recorded. The data must then be recorded in such a way as to ensure that they are reported completely clearly and accurately. The report, whether it be in the form of scientific article, monograph or formal study report, should present the data in a consistent manner and allow for adequate reconstruction of the events which took place during the test. Finally, the report and the data supporting it should be verified to ensure that the test was carried out according to the relevant guidelines (if used), that the study plan was correctly followed and finally that all data were properly generated and accurately reported in the report.(ABSTRACT TRUNCATED AT 250 WORDS)

Databases, Factual↗

Measuring disparities in information capture timeliness across healthcare settings: effects on data quality.

The emergence of evidence-based medicine in the United States has created an industry-wide environment where the quality of data maintained by healthcare organizations is becoming a critical factor in the delivery of medical care. Such a transition necessitates a corresponding need for consistent data collection and maintenance methods. In this study results of a national survey of health information managers were used to assess prevalence of a standard data quality practice, the adoption of policies related to timeliness of data capture. Findings from this survey show that, on a national level, only a slight majority of respondents indicated adoption of timeliness policies. About 61% of respondents indicate they have policies and procedures addressing data timeliness, although persistent patterns of nonadoption were found. We examine how the timeliness of data collection might serve as part of an overall data collection strategy that managers can employ to improve the quality of their information.

Data Collection↗

The RCP Information Laboratory (iLab): breaking the cycle of poor data quality.

A review of data quality in the NHS by the Audit Commission cited a lack of clinician involvement in the validation and use of centrally held activity data as one of the key issues to resolve. The perception that hospital episode statistics cannot support the needs of the individual clinician results in mistrust and disinterest. This in turn leads to under-development of such data from a clinical perspective, and the cycle continues. The RCP Information Laboratory (iLab) aims to address this problem by accessing, analysing and presenting information from these central repositories concerning the activity of visiting individual consultant physicians. With support from iLab staff--an information analyst and a clinician--local data quality issues are highlighted and local solutions sought. The information obtained can be used as an objective measure of activity to support the processes of appraisal and revalidation.

Data Collection↗

Ensuring data quality in medical research through an integrated data management system.

An effective data management system ensures high quality research data by making certain of the proper execution of the study design. This paper presents the components of a data management system and describes procedures for use in each component of the system to obtain high quality data. We discuss the interrelationship among the components of the data management system and the relationship of the data management system to other parts of the research project. We identify underlying principles in design and implementation of a data management system to ensure high quality data.

Computers↗

Managing data quality through automation.

Traditional definitions of data quality deal primarily with individual data sets and the data collection process. Today's standards for ensuring data quality have not changed with respect to the desired results, but have simply been expanded to take advantage of modern technology. Computers are used to acquire, review, store, analyze, and report data. Because each of these steps can be automated, the need for human intervention and manual review is minimized. As a result, the potential for invalid data to reach the data analysis stage has increased significantly. To reduce this potential, efforts must be devoted to developing automated procedures that cover every conceivable validation possibility. Relationships between data and data sets must be well defined [1], and data base support that facilitates ready access to the data for the purpose of analysis must be provided. For small data sets, automation may therefore be impractical; but for large, interrelated data sets, automation is highly desirable. Computer automation has therefore expanded the traditional concept of ensuring data quality to include a complex array of interrelated tasks that must be properly managed to achieve the desired results.

Animals↗

Use of a three-color cDNA microarray platform to measure and control support-bound probe for improved data quality and reproducibility.

Construction methodologies for cDNA microarrays lack the ability to determine array integrity prior to hybridization, leaving the array itself a source of uncontrolled experimental variation. We solved this problem through development of a three-color cDNA array platform whereby printed probes are tagged with fluorescein and are compatible with Cy3 and Cy5 target labeling dyes when using confocal laser scanners possessing narrow bandwidths. Here we use this approach to: (i) develop a tracking system to monitor the printing of probe plates at predicted coordinates; (ii) define the quantity of immobilized probe necessary for quality hybridized array data to establish pre-hybridization array selection criteria; (iii) investigate factors that influence probe availability for hybridization; and (iv) explore the feasibility of hybridized data filtering using element fluorescein intensity. A direct and significant relationship (R2 = 0.73, P < 0.001) between pre-hybridization average fluorescein intensity and subsequent hybridized replicate consistency was observed, illustrating that data quality can be improved by selecting arrays that meet defined pre-hybridization criteria. Furthermore, we demonstrate that our three-color approach provides a means to filter spots possessing insufficient bound probe from hybridized data sets to further improve data quality. Collectively, this strategy will improve microarray data and increase its utility as a sensitive screening tool.

Color↗

Data quality improvement in general practice.

BACKGROUND: The importance of routine data generated by GPs has grown extensively in the last decade. These data have found many applications other than patient care. More attention has therefore been given to the issue of data quality. Several systematic reviews have detected ample space for improvement of data quality. A new review was conducted in order to find out which methods of improvement are effective. METHOD: The Medline database was searched using an iteratively composed set of terms and MeSH (Medical Subject Headings) headings. Only papers that focused on explicit attempts at improving data quality of medical records in general practice were included. RESULTS: Twelve studies met the inclusion criteria. No study used patient-based comparison of records with external sources as the method to assess data quality improvement. Ten studies used internal indicators or markers of data quality instead. Attempts at data quality improvement often involve some sort of individualized feedback, and nearly all attempts seem to have some positive effect. Only one of the included studies fulfilled the basic methodological requirements of an intervention study. The most recent studies used a simple before-after design. CONCLUSION: No intervention to improve data quality has been put to a rigorous enough test. We still lack empirical knowledge as to how improvement can be brought about.

Data Collection↗

Data quality probes-exploiting and improving the quality of electronic patient record data and patient care.

Increasing reliance is being placed on electronic medical records to support clinical care and achieve improved quality standards. In order for clinical information systems (CIS) to deliver excellence the data within it needs to be complete, consistent and accurate. This capture of data is critical but forms only part of the procedure in delivering quality health care during the clinician-patient encounter. A number of processes are involved in this encounter, each of which has to be performed flawlessly to deliver a perfect outcome. This paper outlines a method of assessing the quality of these processes involved in healthcare provision and data quality within a CIS. It proposes the principle of Data Quality Probes (DQP) to assess the performance of the whole encounter system. The main feature of this is the generation of a query which clinical knowledge predicts should not retrieve any cases in a system performing flawlessly. Any cases retrieved (which fail the DQP) indicate an error in either data quality or clinical judgment. This approach is applied practically within the paradigm of a UK family practice testing the hypothesis that a series DQPs can provide a valuable method for monitoring both the data accuracy of a CIS and the provision of quality patient care.

Delivery of Health Care↗

[Data quality in the MORBUS Sentinel Project--expiratory wheezing in infants].

Since 1991 the MORBUS project is being conducted to establish and run a sentinel network of 100 general and paediatric practices in three regions of Germany. A number of health conditions have been and will be monitored consecutively with special emphasis on environmentally determined health problems. From March to June, 1991, 1054 contacts of 1- and 2-year-old children with expiratory wheezing were reported. Quality of these event data was assessed by means of internal completeness. Important clinical information was missing in about 10% of all cases without evidence for regional differentiations. Data quality by this criterion was better in first contact than in re-contact cases (9.7% vs 18.1% missing). Questions concerning the parents (allergies, smoking, education) were less frequently answered (up to 24% missing) than questions of obvious medical relevance to the child. Completeness of parental information varied considerably between regions. There was no association between the medical specialty of the doctors and the quality of their data. In a longitudinal view, there was a slightly positive trend over time in the proportion of clinically incomplete case reports at borderline statistical significance (p = 0.057). Apart from these minor findings then, there was an overall good consistency of completeness in the MORBUS data on expiratory wheezing. By optimizing questionnaires and data transmission, it should be possible to increase the data quality even further.

Air Pollutants↗

Management of data quality--development of a computer-mediated guideline.

Appropriate data quality is a crucial issue in the use of electronically available health data. As source data verification (SDV) and feedback are two standard procedures for measuring and improving data quality it would be worthwhile to adapt these procedures to a current level of quality in order to reduce costs in data management. This project aims to develop a guideline for the management of data quality with special emphasis on this adaptation against the backdrop of research networks in Germany, which operate registers and conduct epidemiological studies. The first step in guideline development was a thorough literature review. The literature offers many measurements as candidates for quality indicators, however, systematic assessments and concepts of SDV and feedback are missing. We assigned possible quality indicators to the levels plausibility, organization, and trueness. Each indicator must be operationally defined to allow automatical calculation. The SDV sample size calculation leads to lower numbers for sites providing data of good quality and larger numbers for sites with poor data quality. The guideline's implementation in a software tool combines two cycles, one for the adaptation of recommendations to a given study/register, the other for the improvement of data quality in a PDCA-like approach. The recommendations will address needs common to medical documentation in daily health care, clinical, epidemiological, and observational studies as well as in surveillance data bases and registers. Further work will have to supplement other aspects of data management.

Feedback↗

Data quality in a distributed data processing system: the SHEP Pilot Study.

The Systolic Hypertension in the Elderly Program (SHEP) Pilot was a collaborative clinical trial that distributed to the clinics all data processing tasks except for randomization assignment codes and morbidity and mortality data. The clinics used customized programs to enter and verify data interactively, to maintain their own local master files, and to transmit the data electronically to the Coordinating Center. We measured quality control based on criteria from centralized as well as distributed models: the error rate for baseline forms was 0.5 per 1000 items. Ninety-eight percent of the forms were query-free, and a central reentry of the data in a 5% sample yielded a miskey rate of 2 per 1000 items. The potential problems of distributed data processing are vulnerability of the local master files and the time demands on Coordinating Center programmers for maintaining clinic computer systems. The advantages are the active involvement of clinic staff in their own quality control, the functional accessibility of the clinics to the Coordinating Center in controlling protocol decisions and data monitoring, and the level of accuracy, completeness, and timeliness of the data that can be achieved.

Aged↗

Measuring quality of life in women with endometriosis: tests of data quality, score reliability, response rate and scaling assumptions of the Endometriosis Health Profile Questionnaire.

BACKGROUND: To test the data quality, scaling assumptions and scoring algorithms underlying the Endometriosis Health Profile-30 (EHP-30) questionnaire: a questionnaire developed to measure the health-related quality of life (HRQoL) of women with endometriosis. METHODS: A cross-sectional postal survey to 727 women with surgically confirmed endometriosis recruited from an existing genetic linkage study (OXEGENE), The National Endometriosis Society (NES), UK and the outpatient gynaecology clinics of the Women's Centre, John Radcliffe Hospital, Oxford. Tests of data quality included secondary factor analysis, internal reliability consistency, descriptive statistics of the data, missing data levels, floor and ceiling effects and corrected item to total correlation scores. RESULTS: Six hundred and ten women (83.9%) returned the questionnaire. Secondary factor analysis verified the domain structure of the EHP-30. All 11 dimensions were internally reliable with Cronbach's alpha scores ranging from 0.80 to 0.96. Missing response rates ranged from 0.2 to 1.3%, and all items were found to be most highly correlated with their own (corrected) scale. CONCLUSIONS: Results confirmed the factor structure, scoring and scaling assumptions of the questionnaire. The high rate of data completeness indicated that the EHP-30 was acceptable and understandable to the respondents, thereby verifying its suitability for measuring the HRQoL of women with endometriosis.

Adolescent↗

Data quality of the Drug Abuse Warning Network.

The purpose of this article was to assess the quality of data collected by the Drug Abuse Warning Network (DAWN), which reports drug abuse emergency department visits. The results of quality assurance studies at 36 sites were reviewed and interpreted. Data collection procedures are not consistent among hospitals and, along with personnel, regularly change within a hospital. Trained investigators reabstracted DAWN report forms at 24 sites and determined that only 57.4% of the cases that met DAWN case definition criteria had been reported; one of five cases had been reported at one site. The technique used in 11 (47.8%) of 23 hospitals to screen for potential DAWN cases detected only 36% of the cases found when all medical charts are examined. The investigators found discrepancies between reported and actual cases in 81.3% of the report forms reabstracted, with an average of 2.3 errors per form. Information as to the drug(s) involved was incorrect in 36.3% of the forms. Due to underreporting of drug abuse emergency department visits and poor quality data in DAWN report forms, DAWN estimates of drug activity must be viewed with caution. Furthermore, estimation of trends is risky, due to differences between emergency departments as to reporting systems and changes over time.

Community Networks↗

Cluster analysis of Delhi's ambient air quality data.

The purpose of this study was to study the spatial patterns of ambient air quality in Delhi in the absence of extensive datasets needed for space-time modeling. A spatial classification was attempted on the basis of ambient air quality data of nine years (1998 is latest year for which published data were available) for three criteria pollutants--nitrogen dioxide, sulfur dioxide, and suspended particulate matter. Monitoring stations take 24-hour samples twice a week. Published monthly average concentration data were used in this study. A hierarchical agglomerative algorithm using the average linkage between groups method and the Euclidean distance metric was used. Cluster analysis indicated that till 1998, by and large, two distinct classes existed. The results of cluster analysis prompted an investigation of systematic biases in the monitored data. No statistically significant differences in the mean concentration of all pollutants were observed between stations belonging to different land-use types (residential and industrial). This fact would be useful, if and when the authorities consider modifying the network or expanding it in Delhi. The results also support the recommendation that Delhi have a uniform standard across all areas. This study has provided a methodology for Indian researchers and practitioners to do an exploratory study of spatial patterns of air pollution and data quality issues in Indian cities using the National Ambient Air Quality Monitoring System data.

Air Pollutants↗

Diagnostic process from the data quality point of view.

The spread of electronic use of data in various areas has put importance of data quality to higher level. Data quality has syntactic and semantic component; the syntactic component is relatively easy to achieve if supported by tools (either off-the-shelf or our own), while semantic component requires more research. In many cases such data come from different sources, are distributed across enterprise and are at different quality levels. Special attention needs to be paid to data upon which critical decisions are met, such as medical data for example. The starting point for research is in our case the risk of the medical area. In the paper we will focus on the semantic component of medical data quality.

Computer Security↗

A comparison of response rate, data quality, and cost in the collection of data on sexual history and personal behaviors. Mail survey approaches and in-person interview.

The authors examined differences in rate of response, data quality, and cost between mail approaches and in-person interview in the collection of data on sexual history and personal behaviors. A sample of women from a midwestern United States university (n = 342) was identified from health service medical records as having been seen for a sexually transmitted disease (cases) or a contraceptive visit (controls) during the latter half of 1985. The women were randomly assigned to one of three data collection strategies. A total of 268 subjects (78%) participated. Results indicated no differences in validity by method of data collection or by case-control status but there were significant differences in completeness, cost, and response rates. In-person interviews resulted in more complete data than mail approaches, although all instruments had low proportions of missing data (0.001-0.006). Response rate differences were not found when data collection methodologies were compared (75-82%) but were found in case-control analyses. Cases were consistently less likely to participate and significantly less likely to respond by mail (p less than 0.05). The cost of the in-person interview was approximately four times that of the mail survey for the data collection. Implications of the case-control response rate difference suggest that mail methodologies, although low in cost, may introduce sampling bias in studies of sexually transmitted diseases.

Adolescent↗

Tests of data quality, scaling assumptions, and reliability of the SF-36 in eleven countries: results from the IQOLA Project. International Quality of Life Assessment.

Data from general population samples in 11 countries (n = 1483 to 9151) were used to assess data quality and test the assumptions underlying the construction and scoring of multi-item scales from the SF-36 Health Survey. Across all countries, the rate of item-level missing data generally was low, although slightly higher for items printed in the grid format. In each country, item means generally were clustered as hypothesized within scales. Correlations between items and hypothesized scales were greater than 0.40 with one exception, supporting item internal consistency. Items generally correlated significantly higher with their own scale than with competing scales, supporting item discriminant validity. Scales could be constructed for 93-100% of respondents. Internal consistency reliability of the eight SF-36 scales was above 0.70 for all scales, with two exceptions. Floor effects were low for all except the two role functioning scales; ceiling effects were high for both role functioning scales and also were noteworthy for the Physical Functioning, Bodily Pain, and Social Functioning scales in some countries. These results support the construction and scoring of the SF-36 translations in these 11 countries using the method of summated ratings.

Cross-Cultural Comparison↗