← Back to match
llamactl
Unified management and routing dashboard for llama.cpp, MLX, and vLLM local LLM models
PlatformWebfreeglobal
llamactl is a self-hosted unified management and routing dashboard for local LLM inference across llama.cpp, MLX, and vLLM backends, with a built-in model downloader for Hugging Face models. It supports dynamic multi-model instances with on-demand loading, automatic idle timeout, and LRU eviction, aimed at developers running multiple local LLM backends who want to manage and route them from one place.
Categories
AI infrastructurelocal LLM tooling

