L1 ~32 KB, ~1 ns. L2 ~512 KB, ~3 ns. L3 ~32 MB, ~10 ns. RAM ~80 ns. Model weights bigger than L3 → every read pays RAM cost.