Built a 768GB VRAM server for frontier models, but open-source is falling behind
The brief
A developer built an EPYC-based server with 768GB VRAM across twelve 64GB cards to run large open-source models locally, but found that frontier open-source options like GLM 5.3 are already lagging behind closed models such as GPT-6 Astra.
Key points
- The post raises questions about the practical limits of local frontier inference as closed models continue to advance.
Sources
- redditreddit.com