Method

How a population-level consumer simulation works — and how far to trust it.

One group, one neutral description, distributions estimated directly, repeated, banded and checked. Here is every step.

  1. 01

    Describe the population, not a persona

    Each target segment becomes a group of 1,000 people described by its composition — age, gender, student status, household income, current spending and alternatives — or job titles, company size and budgets for business buyers — in counts and percentages. Official population statistics (labor and firm statistics for business buyers) supply what they can; the rest is estimated by the model and labeled as such.

  2. 02

    One neutral description

    Every simulated customer reads the same product description, extracted verbatim from your brief. It contains no prices and no claims beyond your sources — and it is the same text real respondents would see.

  3. 03

    Estimate the whole demand curve at once

    Each answer covers every tested price at once, plus a price of zero: the model estimates the share of the group who would buy if they were shown each price, as in a randomized-price survey. The willingness-to-pay distribution follows from the same curve; plan choice is asked separately. The current price is never labelled.

  4. 04

    Perception questions, the same way

    Rating-scale questions — how important, useful, easy to use, satisfying or trustworthy the product seems, and how likely people are to try or recommend it — are answered as distributions too: the share of the group giving each answer. The scale is reversed in half of the runs to cancel position bias.

  5. 05

    Strengthened with real answers to similar questions

    Every simulation also sees real answers to similar purchase questions from a large consumer survey, as a reference for how demand usually falls with price — including people who would not buy even for free. In our validation this markedly improved accuracy (see Evidence).

  6. 06

    Repeat, then widen the band

    Each question runs ten times. The spread across runs is real but too narrow on its own — a model can be consistently wrong. So the uncertainty band also includes a method-error term sized from our backtests, equivalent to a sample of 30 respondents.

  7. 07

    A monotone demand curve

    The curve is averaged across runs, combined across segments by weight and fitted to be non-increasing in price.

  8. 08

    Revenue index and test range

    Revenue index = price × demand. The recommended test range is every price within 10% of the peak. If the peak sits at the edge of your price grid, we tell you to extend it.

  9. 09

    Quality checks on every run

    Validity rate, monotonicity, run-to-run spread, segment separation and where revenue peaks are checked and shown with the results.

  10. 10

    Honest labels

    Until it is calibrated with your own human sample, every output says “Raw simulation — not yet calibrated with this project’s human sample”, and segment differences are labeled directional.

  11. 11

    Calibration with a reserved sample (next)

    The second stage of our design: human responses to the same survey from a small reserved group, used with established statistical methods to estimate and correct prediction errors. Until it ships, every result is labelled as not yet calibrated with your sample.

Method

Backtest

Study 1 · metric JSD

A household survey in China

In a backtest on a national household survey in China (30,000+ households), four five-point questions, a single population-level simulation landed as close to the true answer distribution as a random sample of ≈16–62 real respondents (median 48). To beat one population-level simulation with 95% confidence, a real survey needs ≈37–136 respondents (median ≈111). Agent-level simulation with the same model: ≈4.5–38 (median 5).

≈48

respondents, population-level · break-even (median)

≈111

respondents to beat population-level with 95% confidence (median)

≈5

respondents, agent-level · break-even (median)

Survey question

Distance to the true distribution vs. number of real respondents

Random real sample (mean)5–95% of random samplesPopulation-level simulationAgent-level simulationBreak-even sample sizeSample size to win with 95% confidence
0.00010.0010.010.11.01101001k10kReal respondents sampled at random (log scale)Distance to truth (JSD, log)≈ 49 respondents≈ 4.5 respondents

Answer shares: real vs. simulated

Income far above spending

Slightly above

Balanced

Slightly below

Far below

RealPopulation-level (min–max of 10 runs)Agent-level

Study 2 · metric MAE

A willingness-to-pay survey in the U.S.

In a randomized-price willingness-to-pay survey of U.S. adults (2,058 people, each asked whether they would buy at a randomly assigned price), we took 10 everyday products and had the model estimate directly the share who would buy at each of 11 price points, then compared it with the real purchase rates. Agent-level simulation was worth ≈12 real respondents and population-level simulation ≈17; after strengthening the model with real answers to similar questions, population-level simulation was worth ≈38.

≈12

respondents · agent-level

≈18 to win with 95% confidence

≈17

respondents · population-level

≈25 to win with 95% confidence

≈38

respondents · strengthened population-level

≈53 to win with 95% confidence

Purchase-rate error vs. number of real respondents

Random real sample (mean)5–95% of random samplesAgent-level simulationPopulation-level simulationStrengthened population-levelBreak-even sample sizeSample size to win with 95% confidence
0510152025101001kReal respondents sampled at random (log scale)Purchase-rate error (MAE, points)≈ 12 respondents≈ 17 respondents≈ 38 respondents

Average purchase-rate gap (points, lower is better)

Agent-level simulation17.6 pts
Population-level simulation15.5 pts
Strengthened population-level11.2 pts
Same people, asked again later2.5 pts

The last bar is a reference: when the same respondents answer again later, purchase rates shift by about this much — roughly the floor no method can be expected to beat.

Model calls in study 1

Population-level · 40 calls per group10 repeats × 4 questions
Agent-level · 138,416 callsone call per household per question

log₁₀

What it doesn’t show (yet)

In 13 of 16 subgroup cells, simply reusing the overall distribution matched the subgroup as well or better, so we report segment differences as directional. Pricing validation covers 10 everyday products so far; the model’s errors are mostly systematic and don’t average away with more repeats; and the real answers used for strengthening come from a similar survey, so the gain may shrink for very different categories.

Study 1 metric: Jensen–Shannon divergence between simulated and real answer shares. “Worth n” is the sample size at which the average of 1,000 random real samples is as close as the simulation; “95% confidence” is the size at which 95% of random samples are closer. Backtest v1. Study 2 metric: mean absolute difference between simulated and full-sample real purchase rates across 11 price points (MAE, percentage points), averaged over 10 products; “worth n” is computed as above. Agent-level is a per-respondent simulation (one digital twin per person), using the best-performing published configuration. The strengthened simulation draws on real answers to similar questions about products other than the 10 tested.

Questions

Does this replace talking to customers?

No. It is a fast, inexpensive first read that tells you which prices are worth testing with real customers — and lets you explore far more options than a survey could.

Which models do you use?

Frontier language models from Anthropic. The simulation model, its settings and the prompt version are recorded with every run.

How long does a run take?

A typical run — three segments, six prices, ten repeats — is about 60 model calls and finishes in a few minutes.

Why not simulate individual personas?

Agent-level simulation is great for narratives. For the distribution itself — which is what a demand curve is — we found that modeling the population directly is closer to real data and far cheaper. We still ask for reasons separately.

What data do you keep?

Your brief, uploaded files, selections and results, scoped to your account. You can delete a project at any time, and we never train on it.

Can I calibrate with my own data?

Soon. You will be able to upload a small human sample (30–100 responses) to correct the simulation and tighten the band.

Data sources: study 2 uses the public Twin-2K-500 dataset (Toubia, Gui, Peng, Merlau, Li & Chen, 2025, arXiv:2505.17479), licensed under CC BY 4.0; the reference answers the product uses to strengthen simulations come from the same dataset.

Compare decisions before you commit.

Start with a conversation. Your first simulated readout is a few minutes away.