Search PubMed⌕ Search

PubMed · 12718515

A method to combine non-probability sample data with probability sample data in estimating spatial means of environmental variables.

Abstract

In estimating spatial means of environmental variables of a region from data collected by convenience or purposive sampling, validity of the results can be ensured by collecting additional data through probability sampling. The precision of the pi estimator that uses the probability sample can be increased by interpolating the values at the nonprobability sample points to the probability sample points, and using these interpolated values as an auxiliary variable in the difference or regression estimator. These estimators are (approximately) unbiased, even when the nonprobability sample is severely biased such as in preferential samples. The gain in precision compared to the pi estimator in combination with Simple Random Sampling is controlled by the correlation between the target variable and interpolated variable. This correlation is determined by the size (density) and spatial coverage of the nonprobability sample, and the spatial continuity of the target variable. In a case study the average ratio of the variances of the simple regression estimator and pi estimator was 0.68 for preferential samples of size 150 with moderate spatial clustering, and 0.80 for preferential samples of similar size with strong spatial clustering. In the latter case the simple regression estimator was substantially more precise than the simple difference estimator.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

D J Brus, J J de Gruijter. 2003. A method to combine non-probability sample data with probability sample data in estimating spatial means of environmental variables.. https://doi.org/10.1023/a%3A1022618406507

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Probability estimation when some observations are grouped.

This paper considers the use of additional questions for decreasing survey non-response rates and an approach for estimating a probability based on the results obtained. In a survey, the respondents are asked to answer an original question and follow-up questions, where the answers for the follow-up questions are grouped answers for the original question. For example, respondents are asked to provide an exact number of incidents, but in cases of 'Do not know' or 'Refuse' responses, they are subsequently asked to pick an answer from a less specific categorical scale. The new estimator obtains smaller variance asymptotically and does not depend on a distribution family. This method is applied to income questions in a survey regarding injury prevention and behaviours. Another application is survey data on intimate partner violence, where some amendments were applied for incorporating post-stratification weights and for using non-random grouping. For additional illustration, an example of parameter estimation on artificially generated data is presented.

Data Collection↗