Matchboxmatchbox
← Back to match

llama.cpp

Run open-weight LLMs locally on CPUs, GPUs, and Apple Silicon.

Desktopfreeglobal

A C/C++ inference engine for running open-weight large language models locally or on self-hosted servers, compatible with CPUs, GPUs and Apple Silicon via Metal, CUDA, HIP, SYCL, Vulkan and other backends. It addresses the need to run models without cloud APIs—offering quantization and hybrid CPU/GPU offload to fit models on consumer hardware—and is aimed at developers, machine-learning researchers and privacy-conscious self-hosters who prefer a CLI/server engine over a point-and-click GUI.

Categories
AI / Machine LearningDeveloper ToolsSelf-Hosted

Full match profile

Behind the summary, Matchbox keeps a richer profile of llama.cpp - the signals our matcher actually reads to decide when to surface it. It stays private; claim the listing to see and control it.

  • Problem & pain-point mapping
  • Who we surface it to (audience fit)
  • What it's a strong alternative to
  • Trust & credibility signals

Try Matchbox with your own problem

Describe what is not working - we’ll show you whether llama.cpp (or something else) actually fits.