imp
Self-hosted LLM inference engine built specifically for a single RTX 5090 GPU.
Servicefreeglobal

imp is a free, self-hosted LLM inference engine built from scratch specifically for the NVIDIA RTX 5090, delivering 17-191% faster decoding than llama.cpp on the same models by targeting that GPU's architecture directly. It's designed for agent workloads with tool calls, long context, and many concurrent streams. It's aimed at people running local LLM inference on an RTX 5090 who want maximum throughput from their hardware.
Categories
ai-infrastructureself-hosteddeveloper-tools
Something wrong with this listing — dead link, not a real product, wrong info?

