// ENGINEERING THE FUTURE

High Performance
GPU Kernels.
Built from the Metal Up.

I build and optimize CUDA kernels, profile at scale, and engineer ML infrastructure that delivers measurable performance.

BEST GEMM PERFORMANCE
932.9
GFLOPS on Tesla T4
GPU Exploded Diagram
SCROLL
// FEATURED PROJECTS

cuda-ml-accelerator

Tiled shared-memory GEMM kernel (TILE_WIDTH=32, float32) with Nsight-driven tile size optimization, pybind11 Python bindings, and full benchmark validation against CPU, naive CUDA, and PyTorch/cuBLAS across five matrix sizes.

CUDA Nsight Compute pybind11 float32
VIEW PROJECT →

cuda-kernels

The first CUDA project of the summer — vector addition, naive GEMM, and tiled GEMM (double precision), profiled with Nsight Compute. Found the FP64 pipeline bottleneck that directly motivated the float32 rebuild in cuda-ml-accelerator.

CUDA Nsight Compute double precision
VIEW PROJECT →

CUDA Numerical Integration

Trapezoidal rule and Simpson's rule implemented with CUDA parallel reduction, built alongside Calc I coursework this fall.

CUDA Parallel Reduction Numerical Methods
Planned — September 2026

GPU Kernel Library

A library of high-performance ML kernels — Flash Attention variant, Softmax, LayerNorm, and multi-GPU GEMM — benchmarked against cuDNN.

CUDA Attention LayerNorm Multi-GPU
Planned — Summer 2027