← Back to match
Olla
High-performance proxy and load balancer for LLM infrastructure, unifying local and remote inference backends.
Servicefreeglobal
Olla is a free, open-source, high-performance proxy and load balancer for LLM infrastructure that unifies local and remote inference backends like llama.cpp, Ollama, LM Studio, and vLLM behind intelligent routing with automatic failover. It's aimed at teams running several LLM inference backends who want unified model discovery and reliability instead of hardcoding a single backend into their applications.
Categories
developer toolsai infrastructure

