← Back to match
Shoehorn
Quantizes a BF16 GGUF model to exactly fit your Mac's VRAM, then runs it with llama.cpp.
Desktopfreeglobal
Shoehorn quantizes a BF16 GGUF language model to exactly fit the memory an Apple Silicon Mac actually has, solving a per-tensor mixed-precision assignment guided by an importance matrix instead of using fixed presets. One command downloads, quantizes, and serves the model with llama.cpp. It is open source and macOS-only.
Categories
ailocal llmdeveloper tools

