HomePeopleCompaniesAI ModelsOpen SourceAgentsResearchApps
AllModelsInferenceToolingImage and Video
Open SourceToolingAnthropic 10 Sep 2026 reddit

GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra

The brief

GLM 5.3 Flash at Q4 quantization runs at roughly 60 tokens per second prefill and 550 tokens per second decode on an M3 Ultra Mac, using a Claude Code harness.

Key points

  1. The benchmark makes the model practical for interactive use on Apple Silicon hardware.
Read the original

Sources