vLLM
High-throughput open-source inference engine using paged attention for GPU memory.
why this verdict
Keep it — OpenAI has not replaced it
OpenAI's version overlaps, but does not finish the dev tools job, so this one is still worth keeping open.
- A named launch, not a vibe named, unlinked
- Managed inference APIs removing the need for anyone to operate a serving stack. OpenAI, September 25, 2025 — no announcement link recorded yet.
- How much of the job it covers not the job editorial call
- Parts of it. The job still needs the tool to get finished.
- Is there a free way to do it? yes
- 3 of 3 listed replacements have a usable free tier: SGLang, TensorRT-LLM and Ollama.
- What the call is worth nothing to cancel
- No paid entry tier tracked, so there is no subscription to cancel.
- Threatened by
- OpenAI
- Since
- September 25, 2025
- List price
- free
- Per year
- —
The backstory
vLLM introduced paged attention, which manages GPU memory the way an operating system manages RAM, and multiplied the throughput of self-hosted model serving overnight. It is now the engine under a large share of the inference providers in this archive, which is a strange kind of victory: the thing most exposed to hosted APIs is also what those hosts run. As long as anyone serves open weights at scale, they are almost certainly serving them with this.
Escape hatches
Competing high-performance serving runtime
sglang.ai open_in_newNvidia's optimised serving stack
github.com open_in_newFar simpler for single-machine local use
ollama.com open_in_new3 of 3 replacements have a usable free tier.