← Back to match
Shimmy
Pure-Rust WebGPU inference engine, OpenAI-API compatible, running on any GPU.
Servicefreeglobal
Shimmy is a self-hosted, pure-Rust WebGPU inference engine that is OpenAI-API compatible and GGUF native, running on any GPU as a single binary with no Python or llama.cpp dependency. It's built for developers who want a lightweight, dependency-free local LLM inference server instead of setting up a Python-based inference stack.
Categories
Developer ToolsMachine Learning

