dylan lim
email
twitter
linkedin
hi! i'm a bs/ms student at stanford studying computer science.
i work on ml systems. some recent work includes whole model kernel fusion (
low latency
,
high throughput
) and writing performant
multi-gpu communication primitives
.
prior work
to infinity and beyond, thunderkittens now on nvidia vera rubin
(together ai 2026)
we bought the whole gpu, so we're going to use the whole gpu
(hazy research 2025)
look ma, no bubbles! designing low-latency megakernels
(hazy research 2025)
one kernel for all your gpus
(hazy research 2025)
layoutvlm: 3d scene generation
(cvpr 2025)
flexflow: distributed dnn compiler
(stanford 2024-25)
things i built
shard — distributed training system for local compute
flexgraph — ml model architecture generation for compiler evaluation
industry
together ai — megakernels for everyone
(winter 2025)
jump trading — core strategies
(summer 2025)
valuenex — 10× faster graph algorithms
(spring 2024)
candid — rag pipelines & ml infrastructure
(2023-24)