SGLang
Fast serving runtime for language and vision models with structured output support.
why this verdict
Keep it — OpenAI has not replaced it
OpenAI's version overlaps, but does not finish the dev tools job, so this one is still worth keeping open.
- A named launch, not a vibe named, unlinked
- VLLM's larger community and mindshare in the open-source serving ecosystem. OpenAI, September 25, 2025 — no announcement link recorded yet.
- How much of the job it covers not the job editorial call
- Parts of it. The job still needs the tool to get finished.
- Is there a free way to do it? yes
- 3 of 3 listed replacements have a usable free tier: vLLM, TensorRT-LLM and llama.cpp.
- What the call is worth nothing to cancel
- No paid entry tier tracked, so there is no subscription to cancel.
- Threatened by
- OpenAI
- Since
- September 25, 2025
- List price
- free
- Per year
- —
The backstory
SGLang competes with vLLM on raw serving performance, with a prefix cache design that pays off heavily for agent workloads where many requests share a long system prompt. It also has strong constrained decoding, which matters when outputs must be valid JSON every time. Being second in an open-source category is usually fatal, but inference serving is measured in benchmarks that get rerun constantly, so leadership genuinely changes hands and adoption follows the numbers.
Escape hatches
Larger community and broader hardware support
vllm.ai open_in_newBest raw performance on Nvidia hardware
github.com open_in_newRuns on CPUs and consumer hardware
github.com open_in_new3 of 3 replacements have a usable free tier.