- 3.5× peak speedup on A100
- 6 compiler stages, independently tested
- MXINT8 optional weight quantisation
Elementwise chains in a deep learning model are memory-bound. Each operation reads its input from HBM and writes the result straight back, so a chain of five ops costs five round trips where one would do. The arithmetic is nearly free. The data movement is the bill.
AutoFuser finds those chains and replaces them with a single generated Triton kernel. Same
idea as torch.compile, built from scratch.
Pipeline
- Graph extraction. Trace the model, build the data-dependency graph.
- Fusion detection. Find straight-line chains that can legally collapse into one kernel, respecting dependencies and aliasing.
- Tiling. Choose a tiling scheme per candidate from the tensor shapes.
- Codegen. Emit the fused Triton kernel.
- Auto-tuning. Sweep launch configurations and block sizes, keep the fastest.
- Rewrite. Swap the tuned kernel into the graph in place.
Each stage is a separate module with its own unit tests and benchmarks, so a regression in tiling surfaces as a tiling failure rather than a slowdown three stages downstream.
Results
Up to 3.5× speedup on an NVIDIA A100 across a range of models, with measurable memory-bandwidth reduction. On some models it beats PyTorch’s own TorchInductor.
There is an optional MXINT8 quantisation path for heavy weights, for models that can
absorb the precision loss.