The problem the design was solving
Reasoning models produce long chains of intermediate steps before answering, and score far better than their predecessors on mathematics and coding benchmarks. Whether that constitutes reasoning, or very good pattern matching over problems resembling training data, is the question everyone wanted settled.
Standard benchmarks cannot settle it, for a reason that has nothing to do with model quality. Their problems were published, so a model may have seen them. A high score is consistent with reasoning and with recall, and the benchmark cannot distinguish which.
There is a second problem. Benchmark problems vary in difficulty in ways that are hard to characterise. One question is harder than another, and saying why, or by how much, is not usually possible. So you can observe that a model does worse on a harder set without knowing what dimension of harder is responsible.
A team at Apple set out to fix both, and the design they arrived at is genuinely good. It is worth understanding on its own terms before the criticism, because the criticism is about execution rather than about the idea.

