omp-ninfer
Durable local LLM inference server for coding agents, restart-resumable on one consumer GPU.
Servicefree

omp-ninfer runs a local coding-agent language model on a single consumer GPU with state that survives a server restart, instead of losing in-progress work. It's for developers running local LLM inference for a coding agent who want restart resilience without a cloud API.
Categories
AI infrastructureDeveloper tools
Something wrong with this listing — dead link, not a real product, wrong info?

