← Back to match
tightwad
Pools mismatched GPUs across machines and speeds up local LLM inference with speculative decoding.
PlatformWebfree
Tightwad is a mixed-vendor GPU inference cluster manager that pools CUDA and ROCm GPUs across machines and speeds up local LLM inference with a speculative-decoding proxy: a small model drafts tokens and a larger model verifies them in batch. It plugs into an existing chat app with a single URL change and requires no other workflow adjustment.
Categories
homelablocal AI infrastructure

