GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra
The brief
GLM 5.3 Flash at Q4 quantization runs at roughly 60 tokens per second prefill and 550 tokens per second decode on an M3 Ultra Mac, using a Claude Code harness.
Key points
- The benchmark makes the model practical for interactive use on Apple Silicon hardware.
Sources
- redditreddit.com