Matchboxmatchbox

imp

Self-hosted LLM inference engine built specifically for a single RTX 5090 GPU.

Servicefreeglobal
imp preview

imp is a free, self-hosted LLM inference engine built from scratch specifically for the NVIDIA RTX 5090, delivering 17-191% faster decoding than llama.cpp on the same models by targeting that GPU's architecture directly. It's designed for agent workloads with tool calls, long context, and many concurrent streams. It's aimed at people running local LLM inference on an RTX 5090 who want maximum throughput from their hardware.

Categories
ai-infrastructureself-hosteddeveloper-tools

Full match profile

Behind the summary, Matchbox keeps a richer profile of imp - the signals our matcher actually reads to decide when to surface it. It stays private; claim the listing to see and control it.

  • Problem & pain-point mapping
  • Who we surface it to (audience fit)
  • What it's a strong alternative to
  • Trust & credibility signals

Something wrong with this listing — dead link, not a real product, wrong info?

Try Matchbox with your own problem

Describe what is not working - we’ll show you whether imp (or something else) actually fits.