Shoehorn
Quantizes a BF16 GGUF model to exactly fit your Mac's VRAM, then runs it with llama.cpp.
Desktopfreeglobal

Shoehorn quantizes a BF16 GGUF language model to exactly fit the memory an Apple Silicon Mac actually has, solving a per-tensor mixed-precision assignment guided by an importance matrix instead of using fixed presets. One command downloads, quantizes, and serves the model with llama.cpp. It is open source and macOS-only.
Categories
ailocal llmdeveloper tools
Something wrong with this listing — dead link, not a real product, wrong info?

