Hello, AEA365 community! Liz DiLuzio here, Lead Curator of the blog. This week is Individuals Week, which means we take a break from our themed weeks and spotlight the Hot Tips, Cool Tricks, Rad Resources and Lessons Learned from any evaluator interested in sharing. Would you like to contribute to future individuals weeks? Email me at AEA365@eval.org with an idea or a draft and we will make it happen.
I’m Gene Shackman, Applied Sociologist and author of the Beginners Guide to Evaluation, and I’m Elizabeth DiLuzio, Lead Curator of AEA365.. Welcome to part two about non-probability sampling. In the first blog, I talked about quota sampling, used when administering a survey. The other methods are used after the survey has already been administered.
Weighting after data collection
A second way to make a non-probability sample look like the population is post-survey weighting. As Freese and Jin (2025) put it, participants are assigned different weights after data collection to match known population totals on different characteristics. If you have population data on a variable, you can weight your sample to match the population on that variable.
The American Cancer Society used this approach in a recent multi-state tobacco study. The team surveyed 12,300 adults across 17 states in two waves in 2023 and 2024, drawing the sample from a national opt-in opinion panel whose members signed up through email or online marketing. The researchers weighted each state sample to match Census demographics on age, gender, race and ethnicity, education, marital status, and household size, along with smoker status. They reported that tobacco use trends in the weighted sample tracked national probability surveys, and that demographic distributions matched expected population estimates.
That study used straightforward demographic weighting, but the broader literature describes a wider menu. Arletti, Tanturri, and Paccagnella (2025) review the main approaches.
Raking adjusts weights iteratively until the sample matches the population on each weighting variable’s marginal total. It is simple and works when you only have marginal population totals, but it does not capture interactions between weighting variables.
Propensity score adjustment pairs a non-probability sample with a probability sample, estimates the probability that each observation belongs to the non-probability sample, and uses that probability to reweight. Pollard, Robbins, and Griswold (2026) demonstrate the approach with a Twitter sample and an AmeriSpeak probability sample. The unweighted Twitter sample differed from the probability sample on 24 of 27 political attitudes. After propensity score weighting on demographics, technology use, and political ideology, only 2 of 27 differences remained statistically significant. Political ideology did most of the work; excluding it left 14 attitudes still significantly different. Methodologists are also pushing the framework further. Liang and Wu (2026) propose a survey-weighted propensity score weighting framework for causal inference that combines inverse propensity weights with survey weights to address confounding and selection bias at the same time.
Modeling approaches use the non-probability sample to build a model, then apply the model to pre-dict outcomes for the rest of the population. Multilevel regression and post-stratification (MRP) is the most common version. MRP fits a hierarchical model on the sample, then weights predictions by the size of each demographic cell in the population. The hierarchical structure lets the model borrow information across similar cells, which helps when some cells contain few respondents.
Other approaches include statistical matching, which pairs each probability sample observation with its closest non-probability match, and doubly-robust estimation, which combines propensity scoring with modeling so the estimate stays valid as long as one of the two components is correctly specified.
Across all these methods, the choice of variables matters more than the choice of method. Both Freese and Jin (2025) and Arletti and colleagues reach the same conclusion: picking the right weighting variables drives accuracy more than the technique. None of these methods can correct for selection on a variable you did not measure.
Where the field is heading
Research on non-probability sampling is moving fast. Recent reviews of bias-reduction methods are available here and here. Several journals focus on survey methodology, and a useful list is here.
Non-probability sampling is here to stay. The cost and speed advantages over probability samples are too large to ignore, and the methods for adjusting them are improving. But probability samples are not going away either. As Wu argued in a January 2026 Survey Statistician debate, probability samples or census data remain the most crucial ingredient of any defensible statistical analysis of non-probability samples. You need a benchmark to know whether your adjustments worked.
Hot Tip
Gene curates a website of free resources for social research methods. My page on survey sampling has a section on non-probability sampling.
Do you have questions, concerns, kudos, or content to extend this AEA365 contribution? Please add them in the comments section for this post so that we may enrich our community of practice. Would you like to submit an aea365 Tip? Please send a note of interest to AEA365@eval.org. AEA365 is sponsored by the American Evaluation Association and provides a Tip-a-Day by and for evaluators. The views and opinions expressed on the AEA365 blog are solely those of the original authors and other contributors. These views and opinions do not necessarily represent those of the American Evaluation Association, and/or any/all contributors to this site.
