Matchboxmatchbox
← Back to match

LLMKube

Kubernetes operator for running self-hosted LLM inference across GPU fleets.

PlatformWebfree

LLMKube is a Kubernetes operator for running self-hosted LLM inference across a mixed fleet of NVIDIA, AMD, and Apple Silicon GPUs, using runtimes like llama.cpp and vLLM. It handles multi-GPU sharding, model caching, and exposes OpenAI-compatible endpoints from a declarative spec. It's aimed at platform engineers and homelab or on-prem operators who want to run LLMs on their own hardware without building a custom orchestration platform.

Categories
AI InfrastructureDevOps

Full match profile

Behind the summary, Matchbox keeps a richer profile of LLMKube - the signals our matcher actually reads to decide when to surface it. It stays private; claim the listing to see and control it.

  • Problem & pain-point mapping
  • Who we surface it to (audience fit)
  • What it's a strong alternative to
  • Trust & credibility signals

Try Matchbox with your own problem

Describe what is not working - we’ll show you whether LLMKube (or something else) actually fits.