The cost argument, and why it is the weakest one
Start with the number everyone starts with. Suppose you need 50,000 instruction-response pairs.
Human annotation at two dollars per example, which is modest for anything requiring expertise, costs 100,000 dollars and takes months of coordination.
Generating the same volume at 800 output tokens each, at an assumed three dollars per million output tokens, costs about 120 dollars and runs overnight. That is a ratio of roughly 830 to one.
The ratio is real and it is the least interesting thing about synthetic data, because cost was rarely the binding constraint. Teams that could not afford annotation usually also could not afford to train.
What changes decisions is that generation reaches places collection cannot. Data for a product that has no users yet. Coverage of failure modes that occur once in ten thousand real interactions. Examples in a format no one has ever written down.
The honest framing is that synthetic data is a coverage tool that happens to be cheap, and treating it as a cheap substitute for real data is where the failures start.

