Search PubMed⌕ Search

PubMed · 3844165

Approaches to cleaning data sets: a technical comment.

Abstract

Since each of the methods described has its costs and benefits, more than one method should be used. The combination of methods helps the researcher to obtain an error-free data set. Ideally, value-to-value verification for all data sets is preferred but the multiple entry method is more efficient for large data sets. Each researcher has a definition of large and small data sets. One definition of a large data set is one containing more than 250 records. Such a definition will not be appropriate for all investigators. Therefore, each researcher should balance the allocation of resources needed for a method against the confidence required for the data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

D Y Barhyte, L D Bacon. Approaches to cleaning data sets: a technical comment.. https://pubmed.ncbi.nlm.nih.gov/3844165/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Probability estimation when some observations are grouped.

This paper considers the use of additional questions for decreasing survey non-response rates and an approach for estimating a probability based on the results obtained. In a survey, the respondents are asked to answer an original question and follow-up questions, where the answers for the follow-up questions are grouped answers for the original question. For example, respondents are asked to provide an exact number of incidents, but in cases of 'Do not know' or 'Refuse' responses, they are subsequently asked to pick an answer from a less specific categorical scale. The new estimator obtains smaller variance asymptotically and does not depend on a distribution family. This method is applied to income questions in a survey regarding injury prevention and behaviours. Another application is survey data on intimate partner violence, where some amendments were applied for incorporating post-stratification weights and for using non-random grouping. For additional illustration, an example of parameter estimation on artificially generated data is presented.

Data Collection↗