Observation is not explanation
Everything so far produces hypotheses. A feature that lights up on legal citations, a head that attends to repeated names: these are observations about what correlates with what.
The gap between correlation and mechanism is where interpretability goes wrong. A component can carry information the model never uses downstream. A feature can activate on a concept while contributing nothing to the output. Probing literature calls this the difference between information being present and being used, and the distinction is not academic. A safety claim built on the former is worthless.
The only way to close the gap is intervention. Change something inside the network, hold everything else fixed, and measure whether the behaviour changes. This is ordinary experimental causality applied to a system where, unusually, the experimenter has perfect control: every activation can be read, replaced, or scaled, and the experiment is exactly reproducible.
That control is mechanistic interpretability's real methodological advantage over interpreting biological brains, and this lesson is about how to use it.

