GFLOPS by matrix size, tiled GEMM kernel (TILE_WIDTH=32, float32, Tesla T4). CPU, naive CUDA, and PyTorch/cuBLAS baselines all cross-validated for identical output correctness at every size.
GFLOPS by matrix size · TILE_WIDTH=32 · reaches ~26% of PyTorch/cuBLAS measured throughput at 4096×4096 (~11.5% of T4 FP32 hardware peak)