← Back to match
RAM Coffers
Self-hosted LLM inference tool that cuts cost using NUMA-aware weight banking on refurbished hardware.
Servicefree
RAM Coffers is a self-hosted LLM inference tool that uses NUMA-aware weight banking to speed up inference on refurbished enterprise hardware, reported at up to 8.8x stock llama.cpp performance, avoiding cloud inference costs entirely. It's aimed at people who want to run LLM inference on their own repurposed hardware instead of paying for cloud APIs.
Categories
ai infrastructure

