2026-08-09
Matmul: Correct, but No Tensor Cores
4th kernel: tiled matmul. Correct (matches torch.matmul within fp16 tolerance), but flat around ~1 TFLOPS regardless of matrix size, while PyTorch's cuBLAS climbs to ~38 TFLOPS.
Flat-no-matter-the-size is the tell for a compute-bound kernel — something wasn't engaging. Tried deeper pipelining (num_stages), tried stripping out boundary masking, neither moved the needle. Grepped the actual compiled PTX: zero mma.sync, HMMA, or wmma instructions. Tensor cores never fired, not even once.
Unlike flash attention, this isn't a hard architecture wall — Turing does have tensor cores. Looks more like current Triton (3.6.0) just isn't generating tensor-core code for the T4's sm_75 target anymore, as the project's focus has moved to newer hardware. Second week running an old T4 has surfaced a real gap between what current Triton ships for and what Colab's free-tier GPU can actually do.
