Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Data quality of general practice electronic health records: the impact of a program of assessments, feedback, and training.

OBJECTIVE: The aim of this study was to investigate the impact of a program of repeated assessments, feedback, and training on the quality of coded clinical data in general practice. DESIGN: A prospective uncontrolled intervention study was conducted in a general practice research network. MEASUREMENTS: Percentage of recorded consultations with a coded problem title and percentage of patients receiving a specific drug (e.g., tamoxifen) who had the relevant morbidity code (e.g., breast cancer) were calculated. Annual period prevalence of 12 selected morbidities was compared with parallel data derived from the fourth National Study of Morbidity Statistics from General Practice (MSGP4). RESULTS: The first two measures showed variation between practices at baseline, but on repeat assessments all practices improved or maintained their levels of coding. The period prevalence figures also were variable, but over time rates increased to levels comparable with, or above, MSGP4 rates. Practices were able to provide time and resources for feedback and training sessions. CONCLUSION: A program of repeated assessments, feedback, and training appears to improve data quality in a range of practices. The program is likely to be generalizable to other practices but needs a trained support team to implement it that has implications for cost and resources.

Computer User Training↗

[Data quality of dengue epidemiological surveillance in Belo Horizonte, Southeastern Brazil].

OBJECTIVE: To evaluate the quality of data from the Brazilian information system for mandatory reporting diseases, for the detection of cases notified as suspected dengue fever and hospitalized in the public and private hospitals associated to the Public Health System. METHODS: The study was carried out in Belo Horizonte, Southeastern Brazil, during the years of 1996 to June 2002. The criterion of evaluation used were those recommended by the Guidelines for Evaluating Public Health Surveillance Systems. As a reference standard, medical charts recorded in the Unified System hospitalized discharge database system were revised and validated. A total of 266 (90%) of 294 medical charts were selected; 230 (86.5%) filled the suspect dengue fever criterion. To verify possible association between underreporting and selected variables, was used the odds ratio, with 95% of confidence interval in a logistic regression model. The sensitivity was defined as the proportion of hospitalized dengue cases registered in both systems. Predictive value positive was calculated as the proportion of confirmed cases and those recorded in the reporting system. RESULTS: Underreporting of suspected dengue fever was of 37% cases during 1997 to 2002, it was five times higher during the first three years (OR=5.93; 95% CI: 2.50-14.04) and eight times higher for patients hospitalized in private hospitals than in the public ones (OR=8.42; 95% CI: 2.26-31.27). Underreport was also associated to cases with no haemorrhagic episodes (OR=2.81; 95% CI: 1.28-6.15) and without dengue-specific laboratory exams in medical charts (OR=4.07; 95% CI:1.00-16.52). Sensitivity was 63% and predictive value positive was 43%. CONCLUSIONS: Cases recorded in the reporting system were those more severe and did not represent the total of cases hospitalized in Unified Health System, thus the case fatality rate may be overestimated. The results indicate the necessity of changes in the evaluated surveillance model and in the implementation of the qualification of the health professionals, mainly those working in the private hospitals associated to Unified Health System.

Adolescent↗

Evaluating census data quality using intensive reinterviews: a comparison of U.S. Census Bureau methods and Rasch methods.

"In this paper, we consider Rasch measurement methodology as an alternative to conventional U.S. Census Bureau methodology for evaluating the quality of data obtained using different measurements of the same characteristic. Rather than assuming, like the Census approach, that respondents have true states corresponding to the categories of a questionnaire item, the Rasch approach assumes that measuring instruments and survey respondents vary continuously on one or more common dimensions called latent traits. Unlike the Census approach, the Rasch approach requires that the properties of measuring instruments be invariant across different subclasses of respondents, generating parameter estimates that are sufficient statistics for elementary discrete sampling models. The different conclusions that can be drawn from the Census and Rasch approaches are illustrated by an evaluation of questionnaire items designed to measure the limitation or prevention of work because of physical or mental disability."

Americas↗

Info-tsunami: surviving the storm with data quality probes.

As a result of the rapid expansion of electronically available clinical knowledge, clinicians are faced with potential information overload (info-tsunami). The use of data quality probes (DQPs) in primary care can encourage clinicians' awareness of, and improvement in, data quality entry over time. DQPs can also highlight areas of potential error or omission as well as good practice, which can impact directly upon the quality of patient care. In this paper, five specific conditions have been subjected to the use of a series of DQPs over a five-year period in order to assess and measure the performance of different initiatives on the quality of data capture and patient care.

Asthma↗

An assessment of the data quality for NHEXAS--Part I: Exposure to metals and volatile organic chemicals in Region 5.

A National Human Exposure Assessment Survey (NHEXAS) was performed in U.S. Environmental Protection Agency (U.S. EPA) Region V, providing population-based exposure distribution data for metals and volatile organic chemicals (VOCs) in personal, indoor, and outdoor air, drinking water, beverages, food, dust, soil, blood, and urine. One of the principal objectives of NHEXAS was the testing of protocols for acquiring multimedia exposure measurements and developing databases for use in exposure models and assessments. Analysis of the data quality is one element in assessing the performance of the collection and analysis protocols used in NHEXAS. In addition, investigators must have data quality information available to guide their analyses of the study data. At the beginning of the program quality assurance (QA) goals were established for precision, accuracy, and method quantification limits. The assessment of data quality was complicated. First, quality control (QC) data were not available for all analytes and media sampled, because some of the QC data, e.g., precision of duplicate sample analysis, could be derived only if the analyte was present in the media sampled in at least four pairs of sample duplicates. Furthermore, several laboratories were responsible for the analysis of the collected samples. Each laboratory provided QC data according to their protocols and standard operating procedures (SOPs). Detection limits were established for each analyte in each sample type. The calculation of the method detection limits (MDLs) was different for each analytical method. The analytical methods for metals had adequate sensitivity for arsenic, lead, and cadmium in most media but not for chromium. The QA goals for arsenic and lead were met for all media except arsenic in dust and lead in air. The analytical methods for VOCs in air, water, and blood were sufficiently sensitive and met the QA goals, with very few exceptions. Accuracy was assessed as recovery from field controls. The results were excellent (> or = 98%) for metals in drinking water and acceptable (> or = 75%) for all VOCs except o-xylene in air. The recovery of VOCs from drinking water was lower, with all analytes except toluene (98%) in the 60-85% recovery range. The recovery of VOCs from drinking water also decreased when comparing holding times of < 8 and > 8 days. Assessment of the precision of sample collection and analysis was based on the percent relative standard deviation (% RSD) between the results for duplicate samples. In general, the number of duplicate samples (i.e., sample pairs) with measurable data were too few to assess the precision for cadmium and chromium in the various media. For arsenic and lead, the precision was excellent for indoor, and outdoor air (< 10% RSD) and, although not meeting QA goals, it was acceptable for arsenic in urine and lead in blood, but showed much higher variability in dust. There were no data available for metals in water and food to assess the precision of collection and analysis.

Adolescent↗

Impact of data quality and model complexity on prediction of pesticide leaching.

Accurate input data for leaching models are expensive and difficult to obtain which may lead to the use of "general" non-site-specific input data. This study investigated the effect of using different quality data on model outputs. Three models of varying complexity, GLEAMS, LEACHM, and HYDRUS-2D, were used to simulate pesticide leaching at a field trial near Hamilton, New Zealand, on an allophanic silt loam using input data of varying quality. Each model was run for four different pesticides (hexazinone, procymidone, picloram and triclopyr); three different sets of pesticide sorption and degradation parameters (i.e., site optimized, laboratory derived, and sourced from the USDA Pesticide Properties Database); and three different sets of soil physical data of varying quality (i.e., site specific, regional database, and particle size distribution data). We found that the selection of site-optimized pesticide sorption (Koc) and degradation parameters (half-life), compared to the use of more general database derived values, had significantly more impact than the quality of the soil input data used, but interestingly also more impact than the choice of the models. Models run with pesticide sorption and degradation parameters derived from observed solute concentrations data provided simulation outputs with goodness-of-fit values closest to optimum, followed by laboratory-derived parameters, with the USDA parameters providing the least accurate simulations. In general, when using pesticide sorption and degradation parameters optimized from site solute concentrations, the more complex models (LEACHM and HYDRUS-2D) were more accurate. However, when using USDA database derived parameters, all models performed about equally.

Adsorption↗

Data quality assurance for thermophysical property databases--applications to the TRC SOURCE data system.

To a significant degree processes of database development are based upon human activities, which are susceptible to various errors. Propagation of errors in the processing leads to a decrease in the value of original data as well as that of any database products. Data quality is a critical issue that every database producer must handle as an inseparable part of the database management. Within the Thermodynamics Research Center (TRC), a systematic approach to implement database integrity rules was established through the use of modern database technology, statistical methods, and thermodynamic principles. The four major functions of the system--error prevention, database integrity enforcement, scientific data integrity protection, and database traceability--are detailed in this paper.

Journal Article↗

In-house sulfur SAD phasing: a case study of the effects of data quality and resolution cutoffs.

Single-wavelength anomalous diffraction (SAD) utilizing the weak signal of inherently present S atoms can be successfully used to solve macromolecular structures, although this is mostly performed with data from a synchrotron rather than a laboratory source. Using high redundancy, sufficiently accurate anomalous data may now often be collected in the laboratory using Cu Kalpha X-ray radiation. Systematic analyses of a laboratory-derived data set illuminate the effects of data quality, redundancy and resolution cutoffs on the ability to locate the S atoms and phase the structure of Ptr ToxA, a 13.2 kDa toxin secreted by the fungus Pyrenophora tritici-repentis. Three sulfurs contributed to the successful phasing of the structure and were located using the program SHELXD. It is observed that data quality improves with increasing redundancy, but after a certain point becomes worse owing to crystal decay, so that there is an optimal amount of data to include for the sulfur substructure solution. Further, the success rate in locating S atoms is dramatically improved at lower resolutions and in a manner similar to data quality, there exists an optimal resolution at which the likelihood of solving the substructure is maximized. Based on these observations, a strategy for SAD data collection and substructure solution is suggested.

Crystallization↗

Statistical issues in reporting quality data: small samples and casemix variation.

PURPOSE: To present two key statistical issues that arise in analysis and reporting of quality data. SUMMARY: Casemix variation is relevant to quality reporting when the units being measured have differing distributions of patient characteristics that also affect the quality outcome. When this is the case, adjustment using stratification or regression may be appropriate. Such adjustments may be controversial when the patient characteristic does not have an obvious relationship to the outcome. Stratified reporting poses problems for sample size and reporting format, but may be useful when casemix effects vary across units. Although there are no absolute standards of reliability, high reliabilities (interunit F > or = 10 or reliability > or = 0.9) are desirable for distinguishing above- and below-average units. When small or unequal sample sizes complicate reporting, precision may be improved using indirect estimation techniques that incorporate auxiliary information, and 'shrinkage' estimation can help to summarize the strength of evidence about units with small samples. CONCLUSIONS: With broader understanding of casemix adjustment and methods for analyzing small samples, quality data can be analysed and reported more accurately.

Data Interpretation, Statistical↗

Assessing data quality for decision support--emphasis on secondary analysis.

In secondary analysis, the use of available data makes it possible for the researcher to bypass the most time-consuming and costly steps in the research process. However, there are some noteworthy pitfalls and problems in working with existing data, especially the uncertainty of data quality. If the integrity and quality of data are not assured, statistical analysis of the data will not be reliable, no matter what statistical procedure is used. If the analysis is not reliable, the information used for decision support will not be accurate.

Data Collection↗

HGVbase: a human sequence variation database emphasizing data quality and a broad spectrum of data sources.

HGVbase (Human Genome Variation database; http://hgvbase.cgb.ki.se, formerly known as HGBASE) is an academic effort to provide a high quality and non-redundant database of available genomic variation data of all types, mostly comprising single nucleotide polymorphisms (SNPs). Records include neutral polymorphisms as well as disease-related mutations. Online search tools facilitate data interrogation by sequence similarity and keyword queries, and searching by genome coordinates is now being implemented. Downloads are freely available in XML, Fasta, SRS, SQL and tagged-text file formats. Each entry is presented in the context of its surrounding sequence and many records are related to neighboring human genes and affected features therein. Population allele frequencies are included wherever available. Thorough semi-automated data checking ensures internal consistency and addresses common errors in the source information. To keep pace with recent growth in the field, we have developed tools for fully automated annotation. All variants have been uniquely mapped to the draft genome sequence and are referenced to positions in EMBL/GenBank files. Data utility is enhanced by provision of genotyping assays and functional predictions. Recent data structure extensions allow the capture of haplotype and genotype information, and a new initiative (along with BiSC and HUGO-MDI) aims to create a central repository for the broad collection of clinical mutations and associated disease phenotypes of interest.

Base Sequence↗

Assessment of data quality for the NHEXAS--Part II: Minnesota children's pesticide exposure study (MNCPES).

The Minnesota Children's Pesticide Exposure Study (MNCPES) of the National Human Exposure Assessment Survey (NHEXAS) was conducted in Minnesota to evaluate children's pesticide exposure. This study complements and extends the populations and chemicals included in the NHEXAS Region V study. One of the goals of the study was to test protocols for acquiring exposure measurements and developing databases for use in exposure models and assessments. Analysis of the data quality is one element in assessing the performance of the collection and analysis protocols used in this study. Data quality information must also be available to investigators to guide analysis of the study data. During the planning phase of MNCPES, quality assurance (QA) goals were established for precision, accuracy, and quantification limits. The data quality was assessed against these goals. The assessment is complex. First, data are not available for all analytes and media sampled. In addition, several laboratories were responsible for the analysis of the collected samples. Each laboratory provided data according to their standard operating procedures (SOPs) and protocols. Detection limits were authenticated for each analyte in each sample type. The approach used to calculate detection limits varied across the different analytical methods. The analytical methods for pesticides in air, food, hand rinses, dust wipe and urine were sufficiently sensitive and met the QA goals, with very few exceptions. This was also true for polynuclear aromatic hydrocarbons (PAHs) in air and food. The analytical methods for drinking water and beverages had very low detection limits; however, there were very little measurable data for these samples. The collection and analysis methods for pesticides in surface press samples and soil, and for PAHs in dust wipes were not sufficiently sensitive. Accuracy was assessed primarily as recovery from field controls. The results were good for pesticides and PAHs in air (75-125% recovery). Recovery was lower (<75%) for pesticides in drinking water and beverages. The recovery of pesticides from hand rinses met QA goals (75-100%), but surface press samples showed lower recovery (50-70%). Analysis by gas chromatography-mass spectrometry (GC-MS) did not confirm the presence of atrazine and other pesticides in hand rinse and surface press samples that had been detected by GC-ECD, but instead GC-MS confirmed background interferences. Assessment of the precision of sample collection and analysis is based on the percent relative standard deviation (%RSD) between the results for duplicate samples. Data are available only for pesticides and PAHs in air. Precision was good (<20% RSD) for analytes with measurable data. There were a few analytes with %RSD >20%, but the number of data pairs was very small in these cases. Precision for instrumental analysis of food sample extracts was excellent, with the median %RSD < 20 for all measurable pesticides. The median %RSD for the analysis of replicate aliquots of food from the same sample composite was considerably higher, indicating the potential for inhomogeneity of food homogenates.

Child↗

Data quality assurance and quality control measures in large multicenter stroke trials: the African-American Antiplatelet Stroke Prevention Study experience.

Data quality assurance and quality control are critical to the effective conduct of a clinical trial. In the present commentary, we discuss our experience in a large, multicenter stroke trial. In addition to standard data quality control techniques, we have developed novel methods to enhance the entire process. Central to our methods is the use of clinical monitors who are trained in the techniques of data monitoring.

Journal Article↗

One fish, two fish, we QC fish: controlling data quality among more than 50 organizations over a four-year period.

EPA is conducting a National Study of Chemical Residues in Lake Fish Tissue. The study involves five analytical laboratories, multiple sampling teams from each of the 47 participating states, several tribes, all 10 EPA Regions and several EPA program offices, with input from other federal agencies. To fulfill study objectives, state and tribal sampling teams are voluntarily collecting predator and bottom-dwelling fish from approximately 500 randomly selected lakes over a 4-year period. The fish will be analyzed for more than 300 pollutants. The long-term nature of the study, combined with the large number of participants, created several QA challenges: (1) controlling variability among sampling activities performed by different sampling teams from more than 50 organizations over a 4-year period; (2) controlling variability in lab processes over a 4-year period; (3) generating results that will meet the primary study objectives for use by OW statisticians; (4) generating results that will meet the undefined needs of more than 50 participating organizations; and (5) devising a system for evaluating and defining data quality and for reporting data quality assessments concurrently with the data to ensure that assessment efforts are streamlined and that assessments are consistent among organizations. This paper describes the QA program employed for the study and presents an interim assessment of the program's effectiveness.

Animals↗

Surveying minorities with limited-English proficiency: does data collection method affect data quality among Asian Americans?

BACKGROUND: Little is known about how modes of survey administration affect response rates and data quality among populations with limited-English proficiency (LEP). Asian Americans are a rapidly growing minority group with large numbers of LEP immigrants. OBJECTIVE: We sought to compare the response rates and data quality of interviewer-administered telephone and self-administered mail surveys among LEP Asian Americans. DESIGN: This was a randomized, cross-sectional study using a 78-item survey about quality of medical care that was given to Vietnamese, Mandarin, or Cantonese Chinese patients in their native language. MEASURES: We examined response rates and missing data by mode of survey and language groups. To examine nonresponse bias, we compared the sociodemographic characteristics of respondents and nonrespondents. To assess response patterns, we compared the internal-consistency reliability coefficients across modes and language groups. RESULTS: We achieved an overall response rate of 67% (322 responses of 479 patients surveyed). A higher response rate was achieved by phone interviews (75%) as compared with mail surveys with telephone reminder calls (59%). There were no significant differences in response rates by language group. The mean number of missing item for the mail mode was 4.14 versus 1.67 for the phone mode (P< or =0.000). There were no significant differences in missing data among the language groups and no significant differences in scale reliability coefficients by modes or language groups. CONCLUSIONS: Telephone interviews and mail surveys with phone reminder calls are feasible options to survey LEP Chinese and Vietnamese Americans. These methods may be less costly and labor-intensive ways to include LEP minorities in research.

Adult↗

Census data quality--a user's view.

"This paper presents the perspective of a major user of both decennial and economic [U.S.] census data. It illustrates how these data are used as a framework for commercial marketing research surveys that measure television audiences and sales of consumer goods through retail stores, drawing on Nielsen's own experience in data collection and evaluation. It reviews Nielsen's analyses of census data quality based, in part, on actual field evaluation of census results. Finally, it suggests ways that data quality might be evaluated and improved to enhance the usefulness of these census programs."

Americas↗

Data quality in population-based cancer registration: an assessment of the Merseyside and Cheshire Cancer Registry.

Merseyside and Cheshire Cancer Registry (MCCR) data quality was assessed by applying literature-based measures to 27,942 cases diagnosed in 1990 and 1991. Registrations after death (n = 8535) were also audited (n = 917) to estimate death certificate only (DCO) case accuracy and the proportion of registrations notified by death certificate (DC). Ascertainment appeared to be high from the registration/mortality ratio for lung [1.01:1] and to be low from capture-recapture estimates (59.4%), varying significantly with site from oesophagus [92.2% (95% CI 88.5-95.9)] to breast [47.5 (95% CI 41.8-53.2)]. The estimated DC-dependent proportion was 20% (5601 out of 27 942) with successful traceback in 3533 out of 5601 (63.1%) cases. DCO flagging (2497 out of 27,942, 8.9%) overestimated true DCO cases (2068 out of 27,942, 7.4%). The proportion of cases of unknown primary site was low (1.5%), varying significantly with age [0-4.2%, (95% CI 2.5-5.9)] and district [0.8% (95% CI 0.3-1.3) to 2.2% (95% CI 1.8-2.6)]. The median diagnosis to registration interval appeared to be good (10 weeks), varying significantly with site (P < 0.0001), age (P < 0.0001) and district (P < 0.0001). The proportion with a verified diagnosis was 77.3%, varying significantly with site [lung 55.2% (95% CI 53.7-56.7) to cervix 96.9% (95% CI 96.3-97.5)], age [45.2% (95% CI 40.9-49.5) to 97.5% (95% CI 96.4-98.6)] and district [71.8% (95% CI 69.9-73.8) to 82.5% (95% CI 80.7-84.3)]. The DCO percentages varied similarly by site [non-melanoma skin 0.4% (95% CI 0.2-0.6) to lung 22.6% CI (95% 19.9-25.3)], age [0.7(95% CI 0.1-1.4) to 23.0 (95% CI 19.4-26.6)] and district [6.9% (95% CI 5.7-8.1) to 13.9% (95% CI 12.9-15.0)]. MCCR data quality varied with age, site and district - inviting action - and apparently compares favourably with elsewhere, although deficiencies in published data hampered definitive assessment. Putting quality assurance into practice identified shortcomings in the scope, definition and application of existing measures, and absent standards impeded interpretation. Cancer registry quality assurance should henceforward be within an explicit framework of agreed and standardized measures.

Confidence Intervals↗