Calibrated consumer simulation
Know how customers will respond —
before you commit.
$9.99 a month? a student price? easy enough to use? important to parents?
Tlon simulates your target customers as whole populations and checks its predictions against human evidence, so you can compare decisions before committing resources. Pricing is our first application.
41% would buy at $9.99 a month
IllustrativeThe problem
How will consumers respond — and which decision should we make?
01
4–8 weeks of fieldwork
≈10min
to a first simulated readout
Evidence takes time and budget.
Surveys and market tests need weeks of fieldwork and real budget — and the conditions being studied can change before the research is done.
02
A few options per study
10× repeats
per question and option, so no comparison rests on one answer
Few scenarios get tested.
A new product, a different offer or a price change can each move demand and revenue, yet a real study can only afford to test a few of them.
03
Predictions taken on faith
≈38respondents
what one population-level simulation was worth in our pricing backtest
Simulations need proof.
AI simulation can explore far more scenarios. To act on its results, teams need evidence that the predictions reflect how people actually respond.
The product
From a business question to repeatable simulations.
01
Describe
Tell the agent about your product, who it is for and the decision in front of you. Paste your site, drop a deck. It writes a neutral brief — the exact text every simulated customer reads.
02
Choose who
Pick and adjust the consumers to simulate — students, young adults, parents, or any group you describe. Each group is composed from official population statistics, with every assumption labeled.
03
Simulate
Each group is simulated as a whole population. For a price test the model estimates the whole demand curve in one answer; for perception questions, the share giving each answer. Ten independent runs each, strengthened with real answers to similar questions.
04
Compare
Compare predicted outcomes with honest uncertainty bands: demand and revenue across price points, a recommended test range, and how each group rates the product.
Applications
Pricing first. Perception alongside.More decisions next.
Pricing
Demand and revenue at every tested price, the willingness-to-pay distribution, plan choice and a recommended test range.
Perception & satisfaction
How important, useful, easy to use or satisfying customers expect your product to be, and how likely they are to try or recommend it — by group, with bands.
- Importance44%
- Ease of use61%
- Satisfaction54%
Features & offers
Compare feature sets, bundles and offers before you build them — the same simulations, applied to the other levers of consumer choice.
- Feature set A
- Feature set B
- Bundle
Illustrative
Who to simulate
Tell us who it’s for. We build the crowd.
Young adults, women, students — or any group you can describe. Each composition comes from official US population statistics. Drag the price and see how differently each group responds.
Young adults 50% · Women 68% · Students 22%
50%
would buy at $9.99/mo
- 77.5M US adults
- 29% of all US adults
68%
would buy at $9.99/mo
- 136M US adult women
- 28% of them are aged 18–34
22%
would buy at $9.99/mo
- 21.7M enrolled students
- 56% women · 74% at public colleges
Describe it in a sentence. The agent assembles it from public statistics and labels every assumption.
For example
College students who already pay for music Working moms with young kids Retirees over 65 High-income young professionals
- AgePublic statistic
- SexPublic statistic
- Student statusPublic statistic
- Monthly app spendModel estimate
Method
Simulate the crowd, not the individual.
There are two ways to simulate a market with language models. Both have their place.
Agent-level
Simulate many individual personas one by one, then count their answers.
1,000calls to simulate 1,000 people
Population-level
Describe the whole group once and estimate its answer distribution directly.
1call per group, per question
Unit of simulation
Agent-level ·One persona per call
Population-level ·One group per call
What you get
Agent-level ·Individual answers and narratives
Population-level ·The answer-share distribution
Cost grows with
Agent-level ·Number of simulated people
Population-level ·Questions × repeats
Best at
Agent-level ·Rich stories and per-persona cross-tabs
Population-level ·How answers spread across a population
Watch for
Agent-level ·Answers can bunch on a few options
Population-level ·Subgroup differences come out subtler
We use population-level simulation for the distribution and ask for reasons separately — so you get both the curve and the why.
Calibration
Calibration at the core.
- Stage 1 · Model development
Learn from related surveys
The model draws on human responses to related surveys of similar consumers and products. The backtests below measure this stage.
- Stage 2 · Simulation calibration
Correct with a small reserved sample
After a simulation, human responses to the same survey — same product, questions and test conditions — from a small reserved group are used with established statistical methods to estimate and correct prediction errors.
- Aim
Fewer human responses, same accuracy
The aim is accuracy comparable to a survey with fewer human responses than a survey alone. Independent evaluation will test whether calibration improves accuracy and how much human data it needs.
Longer term, we aim to extend calibration across products and consumer groups, so accumulated consumer data can be reused to explore more scenarios — with fewer repeated studies, and less time and cost.
Design-partner program
We are working with a small cohort of consumer AI companies selling to U.S. consumers. 80% of the research framework is shared; 20% is tailored to your product, plans and competitors.
Tailored · 20%
- 1Baseline
- 2Repeatability
- 3Transfer
- 4Fewer labels
- 5Held-out test
Evidence
One simulation, worth dozens to a hundred-plus real respondents.
Study 1 · metric JSD
A household survey in China
In a backtest on a national household survey in China (30,000+ households), four five-point questions, a single population-level simulation landed as close to the true answer distribution as a random sample of ≈16–62 real respondents (median 48). To beat one population-level simulation with 95% confidence, a real survey needs ≈37–136 respondents (median ≈111). Agent-level simulation with the same model: ≈4.5–38 (median 5).
≈48
respondents, population-level · break-even (median)
≈111
respondents to beat population-level with 95% confidence (median)
≈5
respondents, agent-level · break-even (median)
Survey question
Distance to the true distribution vs. number of real respondents
Answer shares: real vs. simulated
Income far above spending
Slightly above
Balanced
Slightly below
Far below
Study 2 · metric MAE
A willingness-to-pay survey in the U.S.
In a randomized-price willingness-to-pay survey of U.S. adults (2,058 people, each asked whether they would buy at a randomly assigned price), we took 10 everyday products and had the model estimate directly the share who would buy at each of 11 price points, then compared it with the real purchase rates. Agent-level simulation was worth ≈12 real respondents and population-level simulation ≈17; after strengthening the model with real answers to similar questions, population-level simulation was worth ≈38.
≈12
respondents · agent-level
≈18 to win with 95% confidence
≈17
respondents · population-level
≈25 to win with 95% confidence
≈38
respondents · strengthened population-level
≈53 to win with 95% confidence
Purchase-rate error vs. number of real respondents
Average purchase-rate gap (points, lower is better)
The last bar is a reference: when the same respondents answer again later, purchase rates shift by about this much — roughly the floor no method can be expected to beat.
Model calls in study 1
log₁₀
What it doesn’t show (yet)
In 13 of 16 subgroup cells, simply reusing the overall distribution matched the subgroup as well or better, so we report segment differences as directional. Pricing validation covers 10 everyday products so far; the model’s errors are mostly systematic and don’t average away with more repeats; and the real answers used for strengthening come from a similar survey, so the gain may shrink for very different categories.
Study 1 metric: Jensen–Shannon divergence between simulated and real answer shares. “Worth n” is the sample size at which the average of 1,000 random real samples is as close as the simulation; “95% confidence” is the size at which 95% of random samples are closer. Backtest v1. Study 2 metric: mean absolute difference between simulated and full-sample real purchase rates across 11 price points (MAE, percentage points), averaged over 10 products; “worth n” is computed as above. Agent-level is a per-respondent simulation (one digital twin per person), using the best-performing published configuration. The strengthened simulation draws on real answers to similar questions about products other than the 10 tested. Read the full method
What you get
A readout you can act on.
Demand curve
Share who would buy at each price, with a band that includes method error.
Test range · $7.99 – $12.99 /moHover or use ← → to read the curve
Willingness to pay
Distribution across price bands, with median and interquartile range.
p50 $9 · IQR $6–$14
Revenue index
Price × demand. The shaded band is the recommended test range.
Plan mix
Share choosing each plan — or keeping their current setup.
- Basic $4.99
- Plus $9.99
- Pro $19.99
- Keep current setup
Segments
DirectionalPer-segment curves, so you can see who drives the result.
- Young adults
- Women
- Students
Perception
Share giving each answer, lowest to highest; the figure is the share in the top two.
- Importance44%
- Ease of use61%
- Satisfaction54%
Reasons
What each group says for and against buying at the test price — the “why” behind the curve.
- Why they’d buy
- Saves me time every week31%
- Better than the free app I use now22%
- Easy to use on my phone15%
- Why they wouldn’t
- Free alternatives are good enough28%
- Already paying for too many subscriptions19%
- Not sure I’d use it enough13%
Where we start
Pricing for consumer AI companies.U.S. consumers first.
Progress
- DoneA detailed research pipeline, including the calibration design
- DoneAn early demo of the workflow: brief, populations, pricing and perception simulations
- UnderwayModel development
Next milestone
Implement and evaluate the calibration workflow on existing consumer data, starting with pricing: compare calibrated and uncalibrated predictions against independent consumer responses to measure how much calibration improves accuracy.
Security
Your data stays yours.
Project isolation
Every project is scoped to your account with row-level security.
No training on your data
Briefs, files and results are never used to train models.
Delete anytime
Remove a project with its files, runs and reports in one step.
Aggregates first
Populations are built from aggregate statistics, never personal records.
Compare decisions before you commit.
Start with a conversation. Your first simulated readout is a few minutes away.