← Back to match
Reame
CPU-first LLM inference server built to run useful models on cheap ARM and shared-vCPU hardware
PlatformWebfreeglobal
Reame is a CPU-first LLM inference server built on llama.cpp, designed to run usable models efficiently on cheap hardware like shared vCPUs, free-tier cloud boxes, and 2-core ARM machines rather than requiring a GPU. It exposes an OpenAI-compatible API, aimed at developers who want to self-host an LLM on modest CPU-only hardware.
Categories
AI infrastructurelocal LLM tooling

