Fp8 runs ~100 tflops faster when the kernel name has "cutlass" in it

(github.com)

334 points | by mmastrac 2 days ago ago

153 comments