Richael

Blog

2026-08-23

RoPE: One Kernel for Both Directions

6th kernel: rotary position embeddings, the other piece of Kimi K3's attention stack worth building. The nice part isn't the speedup, it's that backward reuses the exact same kernel as forward — just flip the sign on sin.

2026-08-16

RMSNorm: Forward, Backward, and Why It Matters

Reading Kimi K3's report right now, and RMSNorm keeps coming up as one of its bigger architectural wins alongside attention residuals — so this week's kernel is RMSNorm, forward and backward.

2026-08-09

Matmul: Correct, but No Tensor Cores

Output matches PyTorch, but throughput stays flat around 1 TFLOPS no matter the matrix size while cuBLAS climbs to 38 — traced it to the compiled PTX and found zero tensor-core instructions.

2026-08-03

Flash Attention: Correct, but Not Fast

Correctness is solid, but it's ~25x slower than PyTorch on my T4 — turns out the T4 just lacks the pipelining hardware flash attention needs.

2026-07-29

Fused Softmax in Triton

Naive softmax needs 3 passes over memory — fusing them into one pass takes it from ~55 GB/s to ~230 GB/s, a 4x jump, and holds a tighter line than PyTorch's own kernel.

2026-07-26

Writing My First GPU Kernel

Built my first Triton kernel today (vector add), following the official tutorial. Correctness confirmed, and it matches PyTorch's throughput almost exactly on a T4 GPU.

2026-07-25

What Is Soulor AI

Soulor AI is a companion and social simulation platform — not just a chat buddy, but a space to build relationships and rehearse them.

2026-07-25

Hello, world

Starting this site as a place to put my projects and writing.