← Back to match
Modelship
Self-hosted, OpenAI-compatible inference server sharing GPUs across many models through one gateway.
Servicefree
Modelship is a self-hosted, OpenAI-compatible inference server built on Ray Serve, serving reasoning LLMs, universal tool calling, embeddings, speech and image models from one gateway while sharing GPU capacity across them. It's aimed at teams running multiple AI models who want a single self-hosted inference layer instead of separate stacks per model type.
Categories
ai infrastructure

