Three ways to be wrong
Before cataloguing failures, it helps to fix what "working" means, because studies that appear to disagree are often measuring different things.
A silicon sample can be judged on at least three criteria, and they are genuinely independent.
Direction. Does it get the sign right? If human respondents prefer A to B, does the synthetic sample?
Distribution. Does it reproduce the spread of human responses, including disagreement and minority positions, or only the central tendency?
Magnitude. Does it get the size right? A willingness to pay of 12 dollars against a true 40 dollars has the right direction and a useless number.
A method can pass one and fail the others, and the published record contains exactly that pattern. This is why "do silicon samples work?" has no single answer, and why the most common reporting error is demonstrating one criterion and implying all three.

