← Back to match
Expert Streaming Engine
Runs oversized MoE GGUF models via bounded NVMe-RAM-VRAM caching on Linux
Desktopfreeglobal
Expert Streaming Engine is a free, open-source tool for running sparse Mixture-of-Experts GGUF models that don't fit in VRAM, using bounded NVMe-to-RAM-to-VRAM expert caching and native multi-GPU planning, with a Linux desktop Studio, aimed at developers running large local LLMs on limited hardware.
Categories
Developer ToolsLocal AI
Something wrong with this listing — dead link, not a real product, wrong info?

