← Back to match
LLMKube
Kubernetes operator for running self-hosted LLM inference across GPU fleets.
PlatformWebfree
LLMKube is a Kubernetes operator for running self-hosted LLM inference across a mixed fleet of NVIDIA, AMD, and Apple Silicon GPUs, using runtimes like llama.cpp and vLLM. It handles multi-GPU sharding, model caching, and exposes OpenAI-compatible endpoints from a declarative spec. It's aimed at platform engineers and homelab or on-prem operators who want to run LLMs on their own hardware without building a custom orchestration platform.
Categories
AI InfrastructureDevOps

