Blog
2026-08-23
RoPE: One Kernel for Both Directions
6th kernel: rotary position embeddings, the other piece of Kimi K3's attention stack worth building. The nice part isn't the speedup, it's that backward reuses the exact same kernel as forward — just flip the sign on sin.
2026-08-16
RMSNorm: Forward, Backward, and Why It Matters
Reading Kimi K3's report right now, and RMSNorm keeps coming up as one of its bigger architectural wins alongside attention residuals — so this week's kernel is RMSNorm, forward and backward.
2026-08-09
Matmul: Correct, but No Tensor Cores
Output matches PyTorch, but throughput stays flat around 1 TFLOPS no matter the matrix size while cuBLAS climbs to 38 — traced it to the compiled PTX and found zero tensor-core instructions.
2026-08-03
Flash Attention: Correct, but Not Fast
Correctness is solid, but it's ~25x slower than PyTorch on my T4 — turns out the T4 just lacks the pipelining hardware flash attention needs.
2026-07-29
Fused Softmax in Triton
Naive softmax needs 3 passes over memory — fusing them into one pass takes it from ~55 GB/s to ~230 GB/s, a 4x jump, and holds a tighter line than PyTorch's own kernel.
2026-07-26
Writing My First GPU Kernel
Built my first Triton kernel today (vector add), following the official tutorial. Correctness confirmed, and it matches PyTorch's throughput almost exactly on a T4 GPU.
2026-07-25
What Is Soulor AI
Soulor AI is a companion and social simulation platform — not just a chat buddy, but a space to build relationships and rehearse them.
2026-07-25
Hello, world
Starting this site as a place to put my projects and writing.