← Back to match
llama.cpp
Run open-weight LLMs locally on CPUs, GPUs, and Apple Silicon.
Desktopfreeglobal
A C/C++ inference engine for running open-weight large language models locally or on self-hosted servers, compatible with CPUs, GPUs and Apple Silicon via Metal, CUDA, HIP, SYCL, Vulkan and other backends. It addresses the need to run models without cloud APIs—offering quantization and hybrid CPU/GPU offload to fit models on consumer hardware—and is aimed at developers, machine-learning researchers and privacy-conscious self-hosters who prefer a CLI/server engine over a point-and-click GUI.
Categories
AI / Machine LearningDeveloper ToolsSelf-Hosted

