swe-bench
3 free lessons tagged swe-bench across AI, Business. Each one is a short sequence of focused steps with narration and a five-question quiz at the end — take them in any order, no signup required.
Evaluating AI Agents: From Final Answers to Full Trajectories
A rigorous look at how to measure agent performance — trajectory-level vs final-answer evals, canonical multi-step benchmarks (SWE-bench, WebArena, OSWorld, GAIA), LLM-as-judge pitfalls, and why your eval is your spec.
Coding Agents: Productivity Studies, Benchmarks, and the METR Slowdown
Hard evidence on AI coding assistants: the Microsoft Research RCTs, METR's surprising 2025 slowdown finding, SWE-bench Verified, and why benchmark scores keep outrunning real-world value.
Interpreting LLM Benchmarks: What MMLU, GPQA, and SWE-bench Actually Measure
A field guide to the LLM benchmarks practitioners cite in 2026 — what each one measures, where it's saturated, where contamination risk is high, and why benchmark gains rarely transfer to your task without your own eval.

