280 tok/s on Qwen3.8 27B Using Dual AMD R9700s and 940k Token KV Cache
TLDR
A local inference enthusiast documented achieving 280 tokens per second on Qwen3.8 27B using two AMD Radeon RX 9700 GPUs with a 940k token KV cache. The writeup covers driver setup, ROCm configuration, and the specific llama.cpp flags that unlocked the result. Dual R9700 performance figures are scarce, making this a useful reference for the AMD open-source inference path.