Ferrox
Pure-Rust GGUF inference engine with CPU, Metal and CUDA kernels, benchmarked vs llama.cpp.
PlatformWebfreeglobal
Ferrox is a pure-Rust GGUF inference engine supporting dense and Mixture-of-Experts models on CPU, Apple Metal or CUDA, with an OpenAI-compatible server, benchmarked head-to-head against llama.cpp. It targets developers running local LLM inference who want a native Rust engine rather than binding to a C++ library.
Categories
AI InfrastructureDeveloper Tools
Something wrong with this listing — dead link, not a real product, wrong info?

