My Local Model Setup on an M4 Pro Mac Mini
TLDR
A detailed walkthrough of running local LLMs on an M4 Pro Mac Mini, covering model selection, inference tools, and performance tuning. The author documents which models run well within Apple Silicon's unified memory constraints and the practical tradeoffs between speed and quality. The post serves as a reference guide for anyone setting up on-device inference without cloud dependency.