Whallm
Runs large DeepSeek and Qwen models on Apple Silicon Macs by streaming experts from SSD.
Desktopfreeglobal

Whallm runs large DeepSeek and Qwen mixture-of-experts models on Apple Silicon Macs by streaming only the needed experts from SSD instead of requiring the full model in RAM, shipping with a built-in chat UI and an OpenAI-compatible API. It targets Mac users who want to run very large local LLMs without high-RAM hardware.
Categories
AI InfrastructureDeveloper Tools
Something wrong with this listing — dead link, not a real product, wrong info?

