← Back to blog
2026-07-29
Fused Softmax in Triton
Mid-week 2nd kernel: fused softmax. Naive softmax needs 3 passes over memory (max, exp+sum, divide) — fusing into one pass takes it from ~55 GB/s to ~230 GB/s, a 4x jump.
Triton's line is also tighter and more consistent than PyTorch's across block sizes.
Next: flash attention.
