AnyLearn
All lessons
Businessadvanced

Silicon Samples: Surveying a Model Instead of People

Where the idea of using language models as synthetic survey respondents came from, the three founding results that made marketing take it seriously, and the economics that make it so tempting.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 9

The cost structure that invites disruption

Consumer research runs on a slow, expensive input: people. A nationally representative survey means recruiting a panel, screening for quotas, paying incentives, and waiting. A conjoint study measuring how buyers trade off price against features needs hundreds of respondents each answering dozens of tedious comparisons. Turnaround is measured in weeks and budgets in tens of thousands.

Every step of that is a constraint on what gets researched. Small questions do not justify a study, so they get answered by opinion. Iterations are rationed. Segments too small to sample economically go unstudied.

So when a technology arrives that appears to answer survey questions in the voice of any demographic you specify, instantly and at negligible cost, the pull is enormous. That is what a silicon sample is: a synthetic dataset generated by a language model, intended to stand in for human respondents.

The question this path answers is not whether it is tempting. It is whether the data is any good.

Full lesson text

All 9 steps on one page, for reading, reference, and search.

Show

1. The cost structure that invites disruption

Consumer research runs on a slow, expensive input: people. A nationally representative survey means recruiting a panel, screening for quotas, paying incentives, and waiting. A conjoint study measuring how buyers trade off price against features needs hundreds of respondents each answering dozens of tedious comparisons. Turnaround is measured in weeks and budgets in tens of thousands.

Every step of that is a constraint on what gets researched. Small questions do not justify a study, so they get answered by opinion. Iterations are rationed. Segments too small to sample economically go unstudied.

So when a technology arrives that appears to answer survey questions in the voice of any demographic you specify, instantly and at negligible cost, the pull is enormous. That is what a silicon sample is: a synthetic dataset generated by a language model, intended to stand in for human respondents.

The question this path answers is not whether it is tempting. It is whether the data is any good.

2. Why a text predictor might encode people at all

The premise sounds implausible until you look at what training actually does.

A language model is fit to predict text written by humans, at enormous scale. That corpus contains product reviews, forum arguments, survey write-ups, complaints, recommendations, and the ordinary written record of people expressing preferences. To predict that text well, a model has to capture regularities in how different kinds of people talk about different kinds of things.

The claim is not that the model has preferences. It is that the conditional distribution of text given a described person carries information about how such people actually respond, because that conditioning was learned from text those people wrote.

If that holds, then prompting with a demographic description and reading off an answer is a sampling procedure, drawing from a learned approximation of a subpopulation's response distribution.

That is the entire bet, and it is testable. The rest of this lesson covers the three results that made researchers take it seriously.

3. Algorithmic fidelity

The founding paper is "Out of One, Many: Using Language Models to Simulate Human Samples" by Lisa Argyle, Ethan Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, and David Wingate, published in Political Analysis volume 31, issue 3, pages 337 to 351, in 2023. It gave the field both its method and its vocabulary.

The method: take thousands of real socio-demographic backstories from participants in large existing surveys, condition the model on each one, and collect its answers. Because the human answers are known, correspondence can be measured rather than asserted.

The finding was that bias in the model is not a uniform smear. It is fine-grained and demographically correlated, so conditioning on a detailed backstory shifts the output distribution in ways that track the corresponding human subgroup.

They named this property algorithmic fidelity, and named the technique silicon sampling. Both terms are now standard.

Notice what the concept concedes: fidelity is a property a model has to a degree, in a domain. It is not a switch that is on.

4. Homo silicus

Economics arrived at the same idea from a different direction. John Horton's "Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?" (NBER Working Paper 31122, 2023) proposed treating the model as an implicit computational model of a human, a homo silicus, usable the way economists use homo economicus.

The experimental design is the appealing part. Rather than asking the model to predict what people do, Horton re-ran classic behavioural economics experiments on it: endow it, inform it, give it preferences, then observe its choices in the original scenario.

Source experiments included work by Charness and Rabin on social preferences, Kahneman, Knetsch and Thaler on fairness in pricing, and Samuelson and Zeckhauser on status quo bias. The simulated results were qualitatively similar to the originals.

One objection gets addressed directly. These experiments are famous and their results are in the training corpus, so the model might be reciting. Horton reports that it does not simply overfit to known results and generalises to novel scenarios.

5. The marketing result: willingness to pay

The result that landed hardest in marketing came from James Brand, Ayelet Israeli, and Donald Ngwe, in a Harvard Business School working paper (number 23-062, circulated from 2023, titled "Using GPT for Market Research" and later "Using LLMs for Market Research").

They prompted the model to answer survey questions as a shopper in a product category, and estimated willingness to pay from its responses. Three findings mattered because they are structural, not cosmetic.

The model produces a downward-sloping demand curve: query it thousands of times across prices and fewer synthetic shoppers buy as price rises. It shows declining marginal utility for additional units. And it responds to stated income: told the shopper earns more, it becomes less price-sensitive.

Those are not arbitrary outputs. They are the qualitative shape economics predicts, reproduced without being asked for.

The same paper contains the finding this path returns to repeatedly, and it is easy to skip past in the enthusiasm: the willingness-to-pay estimates are sometimes comparable to human studies, but are often inaccurate and in some cases wrong-signed.

6. The silicon sampling pipeline

The workflow mirrors a real study closely enough that each stage has a human counterpart, which is exactly why validity has to be argued stage by stage rather than assumed from the output looking plausible.

flowchart TD
  A["Define the target population"] --> B["Write demographic backstories"]
  B --> C["Condition the model on each backstory"]
  C --> D["Ask the survey questions"]
  D --> E["Collect responses as a dataset"]
  E --> F["Analyse as if it were panel data"]
  F --> G{"Validated against human data?"}
  G -- yes --> H["Findings carry limited weight"]
  G -- no --> I["Findings carry none"]

7. The economics, quantified

It is worth being concrete about the appeal, because the size of the gap explains why adoption is running ahead of validation.

Conjoint study, 800 respondents, 20 choice tasks each

Human panel
  recruitment + incentives      $15,000 - $40,000
  fielding time                 2 - 4 weeks
  iterations affordable         1, maybe 2
  new segment                   full re-field

Silicon sample
  API cost                      tens of dollars
  generation time               hours
  iterations affordable         effectively unlimited
  new segment                   edit the backstory text

The cost ratio is roughly three orders of magnitude, and the time ratio is similar. But the operational difference matters more than either: when a study costs almost nothing, you can run the variant, test the tail segment, and iterate the wording. That changes what questions are askable, not just what they cost.

And it is precisely why the validity question cannot be waved through. A cheap instrument that is wrong does not save money. It produces confident decisions on false evidence, faster and in greater volume than an expensive one could.

8. The necessary condition

Marketing's central treatment is "Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines" by Marko Sarstedt, Susanne Adler, Lea Rau, and Bernd Schmitt, in Psychology & Marketing volume 41, pages 1254 to 1270, 2024.

Their framing is the one to carry forward. Replacing human respondents with synthetic ones requires that the generated data generalises to human respondents. That is a necessary condition, and it is an empirical question with a different answer in every domain.

Reviewing studies comparing silicon and human samples, they report that result patterns vary considerably across domains. So the honest answer to "do silicon samples work?" is not yes or no. It is: for which construct, which population, and which question format, and has anyone checked?

One predictor recurs across this literature. Domains well represented in training text, such as evaluations of familiar products and services, tend to fare better than sparse or novel domains, such as consumer response to an unfolding crisis. Coverage in the corpus is doing the work.

9. What a silicon sample is not

Four category errors account for most misuse, and naming them now prevents them later.

It is not a panel. There are no independent respondents. One model produces every answer, so what looks like 800 people is 800 draws from one conditional distribution. Standard errors computed as though observations were independent are meaningless, and generating more rows shrinks them without adding information.

It is not a discovery instrument. It can only recombine what is in its training distribution. Asking it what people will think of a product category that does not exist yet is asking it to interpolate, and it will comply fluently.

It is not unbiased because it is synthetic. It inherits the composition of its training corpus and its alignment training. Groups underrepresented online are underrepresented here, with no sampling frame to audit.

It is not a substitute for a validation study. Fidelity demonstrated for political attitudes says nothing about fidelity for snack preference.

What it is: a fast, cheap, plausible approximation whose accuracy is unknown until measured. The next lesson measures it.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. What is 'algorithmic fidelity' as Argyle and colleagues defined it?
    • The rate at which a model reproduces its training data verbatim
    • The property that a model's bias is fine-grained and demographically correlated, so conditioning emulates subgroup response distributions
    • The accuracy of a model's factual claims about demographics
    • The consistency of a model's answers when the same prompt is repeated
  2. Which findings did Brand, Israeli, and Ngwe report from prompting a model as a shopper?
    • Perfectly calibrated willingness-to-pay matching human panels across categories
    • Random responses with no economic structure
    • A downward-sloping demand curve and declining marginal utility, but estimates often inaccurate and sometimes wrong-signed
    • Upward-sloping demand indicating the model misunderstands prices
  3. Why are standard errors computed from a silicon sample as though respondents were independent misleading?
    • Language models produce responses with unusually high variance
    • The sample size is always too small for asymptotic approximations
    • Synthetic responses require a different distributional family
    • Every response comes from one model, so extra rows shrink standard errors without adding information
  4. According to Sarstedt and colleagues, what is the necessary condition for replacing human samples with silicon ones?
    • That the generated data generalizes to human respondents, which varies considerably by domain
    • That the model be fine-tuned on the target product category
    • That the sample size exceed the equivalent human study
    • That responses be generated by more than one model provider
  5. What predicts whether a silicon sample performs relatively well in a given domain?
    • The complexity of the statistical model applied afterward
    • How well represented that domain is in the model's training text
    • The number of demographic variables in the backstory
    • Whether the questions are open-ended rather than closed

Related lessons

AI
intermediate

What This Teaches About Measuring Anything

The exchange is a case study with transferable rules. A conclusion resting on failures needs a failure taxonomy. Every instance must be verified solvable before anyone is scored against it. Output format is a confound whenever answers get long. And when two explanations fit the same data, the productive move is to find the prediction on which they differ, then test it.

8 steps·~12 min
AI
intermediate

The Rebuttal: Three Ways to Score Zero Without Failing

The response disputed none of the data and argued the experiment measured something other than reasoning. Models had to print move lists exceeding their output limits, and said so in the transcripts. Some instances had no solution and were scored as failures anyway. And asking for a program instead of a move list produced high accuracy on instances reported as total collapse.

8 steps·~12 min
AI
intermediate

The Experiment: Puzzles With a Difficulty Dial

Apple researchers built an evaluation designed to fix a real problem with benchmarks: puzzles where difficulty turns up smoothly while the logic stays identical, and every step can be checked. They found accuracy collapsing to zero past a threshold, and not improving when the solution algorithm was handed to the model. This lesson covers the design and why it was a good one.

8 steps·~12 min
AI
intermediate

Ten Domains, and a Profile That Is Not Flat

The framework scores ten cognitive domains at ten percent each. Running it produces something more useful than the headline totals of 27 percent for GPT-4 and around 57 for GPT-5: a jagged profile, where a model is at or near full marks on some domains and at zero on others. This lesson walks the domains, reads both profiles column by column, and shows what the jaggedness explains.

8 steps·~12 min