AnyLearn
All cursus
AIadvanced

Mechanistic Interpretability: Reverse-Engineering Neural Networks

Attribution methods tell you which inputs mattered. Mechanistic interpretability asks the harder question: what algorithm is the network actually running? This path builds the field from its foundations, why individual neurons are the wrong unit and what superposition says is really happening, through the circuits that implement in-context learning, the sparse autoencoders that pulled 34 million features out of a production model, and the causal interventions that separate a tested claim from a plausible story.

0 of 4 lessons complete
Sign in to track progress and earn a certificate.

Lessons, in order

  1. 1
    AI
    Features, Directions, and Superposition
    Start
  2. 2
    AI
    Circuits: Reverse-Engineering Transformer Algorithms
    Start
  3. 3
    AI
    Sparse Autoencoders and Dictionary Learning
    Start
  4. 4
    AI
    Causal Interventions, Steering, and What Remains Open
    Start