Method
How a population-level consumer simulation works — and how far to trust it.
One group, one neutral description, distributions estimated directly, repeated, banded and checked. Here is every step.
01
Describe the population, not a persona
Each target segment becomes a group of 1,000 people described by its composition — age, gender, student status, household income, current spending and alternatives — or job titles, company size and budgets for business buyers — in counts and percentages. Official population statistics (labor and firm statistics for business buyers) supply what they can; the rest is estimated by the model and labeled as such.
02
One neutral description
Every simulated customer reads the same product description, extracted verbatim from your brief. It contains no prices and no claims beyond your sources — and it is the same text real respondents would see.
03
Estimate the whole demand curve at once
Each answer covers every tested price at once, plus a price of zero: the model estimates the share of the group who would buy if they were shown each price, as in a randomized-price survey. The willingness-to-pay distribution follows from the same curve; plan choice is asked separately. The current price is never labelled.
04
Perception questions, the same way
Rating-scale questions — how important, useful, easy to use, satisfying or trustworthy the product seems, and how likely people are to try or recommend it — are answered as distributions too: the share of the group giving each answer. The scale is reversed in half of the runs to cancel position bias.
05
Strengthened with real answers to similar questions
Every simulation also sees real answers to similar purchase questions from a large consumer survey, as a reference for how demand usually falls with price — including people who would not buy even for free. In our validation this markedly improved accuracy (see Evidence).
06
Repeat, then widen the band
Each question runs ten times. The spread across runs is real but too narrow on its own — a model can be consistently wrong. So the uncertainty band also includes a method-error term sized from our backtests, equivalent to a sample of 30 respondents.
07
A monotone demand curve
The curve is averaged across runs, combined across segments by weight and fitted to be non-increasing in price.
08
Revenue index and test range
Revenue index = price × demand. The recommended test range is every price within 10% of the peak. If the peak sits at the edge of your price grid, we tell you to extend it.
09
Quality checks on every run
Validity rate, monotonicity, run-to-run spread, segment separation and where revenue peaks are checked and shown with the results.
10
Honest labels
Until it is calibrated with your own human sample, every output says “Raw simulation — not yet calibrated with this project’s human sample”, and segment differences are labeled directional.
11
Calibration with a reserved sample (next)
The second stage of our design: human responses to the same survey from a small reserved group, used with established statistical methods to estimate and correct prediction errors. Until it ships, every result is labelled as not yet calibrated with your sample.
Method
Backtest
Study 1 · metric JSD
A household survey in China
In a backtest on a national household survey in China (30,000+ households), four five-point questions, a single population-level simulation landed as close to the true answer distribution as a random sample of ≈16–62 real respondents (median 48). To beat one population-level simulation with 95% confidence, a real survey needs ≈37–136 respondents (median ≈111). Agent-level simulation with the same model: ≈4.5–38 (median 5).
≈48
respondents, population-level · break-even (median)
≈111
respondents to beat population-level with 95% confidence (median)
≈5
respondents, agent-level · break-even (median)
Survey question
Distance to the true distribution vs. number of real respondents
Answer shares: real vs. simulated
Income far above spending
Slightly above
Balanced
Slightly below
Far below
Study 2 · metric MAE
A willingness-to-pay survey in the U.S.
In a randomized-price willingness-to-pay survey of U.S. adults (2,058 people, each asked whether they would buy at a randomly assigned price), we took 10 everyday products and had the model estimate directly the share who would buy at each of 11 price points, then compared it with the real purchase rates. Agent-level simulation was worth ≈12 real respondents and population-level simulation ≈17; after strengthening the model with real answers to similar questions, population-level simulation was worth ≈38.
≈12
respondents · agent-level
≈18 to win with 95% confidence
≈17
respondents · population-level
≈25 to win with 95% confidence
≈38
respondents · strengthened population-level
≈53 to win with 95% confidence
Purchase-rate error vs. number of real respondents
Average purchase-rate gap (points, lower is better)
The last bar is a reference: when the same respondents answer again later, purchase rates shift by about this much — roughly the floor no method can be expected to beat.
Model calls in study 1
log₁₀
What it doesn’t show (yet)
In 13 of 16 subgroup cells, simply reusing the overall distribution matched the subgroup as well or better, so we report segment differences as directional. Pricing validation covers 10 everyday products so far; the model’s errors are mostly systematic and don’t average away with more repeats; and the real answers used for strengthening come from a similar survey, so the gain may shrink for very different categories.
Study 1 metric: Jensen–Shannon divergence between simulated and real answer shares. “Worth n” is the sample size at which the average of 1,000 random real samples is as close as the simulation; “95% confidence” is the size at which 95% of random samples are closer. Backtest v1. Study 2 metric: mean absolute difference between simulated and full-sample real purchase rates across 11 price points (MAE, percentage points), averaged over 10 products; “worth n” is computed as above. Agent-level is a per-respondent simulation (one digital twin per person), using the best-performing published configuration. The strengthened simulation draws on real answers to similar questions about products other than the 10 tested.
Questions
Does this replace talking to customers?
No. It is a fast, inexpensive first read that tells you which prices are worth testing with real customers — and lets you explore far more options than a survey could.
Which models do you use?
Frontier language models from Anthropic. The simulation model, its settings and the prompt version are recorded with every run.
How long does a run take?
A typical run — three segments, six prices, ten repeats — is about 60 model calls and finishes in a few minutes.
Why not simulate individual personas?
Agent-level simulation is great for narratives. For the distribution itself — which is what a demand curve is — we found that modeling the population directly is closer to real data and far cheaper. We still ask for reasons separately.
What data do you keep?
Your brief, uploaded files, selections and results, scoped to your account. You can delete a project at any time, and we never train on it.
Can I calibrate with my own data?
Soon. You will be able to upload a small human sample (30–100 responses) to correct the simulation and tighten the band.
Data sources: study 2 uses the public Twin-2K-500 dataset (Toubia, Gui, Peng, Merlau, Li & Chen, 2025, arXiv:2505.17479), licensed under CC BY 4.0; the reference answers the product uses to strengthen simulations come from the same dataset.
Compare decisions before you commit.
Start with a conversation. Your first simulated readout is a few minutes away.