
vLLM
Open-source inference and serving engine for running large language models at scale.
PlatformWebfreeglobal

vLLM is an open-source inference and serving engine for running large language models on owned or rented GPUs, built for high throughput and efficient memory use rather than managed-API costs. It offers an OpenAI-compatible API and runs on NVIDIA, AMD, Intel, and CPU-only hardware via Python or Docker. It's aimed at developers and teams who want to self-host LLM inference at scale.
Categories
Developer ToolsAI Infrastructure
Something wrong with this listing — dead link, not a real product, wrong info?

