// BENCHMARKS

cuda-ml-accelerator

GFLOPS by matrix size, tiled GEMM kernel (TILE_WIDTH=32, float32, Tesla T4). CPU, naive CUDA, and PyTorch/cuBLAS baselines all cross-validated for identical output correctness at every size.

256²
162.1
512²
448.1
1024²
401.0
2048²
598.6
4096²
932.9

GFLOPS by matrix size · TILE_WIDTH=32 · reaches ~26% of PyTorch/cuBLAS measured throughput at 4096×4096 (~11.5% of T4 FP32 hardware peak)