Skip to content
L llama.cpp logo

llama.cpp

code Dev Tools · free tier
NOT REALLY

C++ inference engine that runs quantised language models on ordinary hardware.

NOT REALLY

why this verdict

Keep it — Meta has not replaced it

Meta's version overlaps, but does not finish the dev tools job, so this one is still worth keeping open.

A named launch, not a vibe named, unlinked
Cheap hosted APIs making local inference a hobby rather than a requirement. Meta, September 25, 2025 — no announcement link recorded yet.
How much of the job it covers not the job editorial call
Parts of it. The job still needs the tool to get finished.
Is there a free way to do it? yes
3 of 3 listed replacements have a usable free tier: Ollama, LM Studio and MLX.
What the call is worth nothing to cancel
No paid entry tier tracked, so there is no subscription to cancel.
Reason composed from the fields above missing: announcement linked recorded: free replacement listed recorded: two or more escape hatches
Threatened by
Meta
Since
September 25, 2025
List price
free
Per year
—

The backstory

llama.cpp proved that a capable model could run on a laptop with no GPU, and its GGUF quantisation format became the standard for distributing models people actually run at home. Hosted APIs are cheaper than most people's time, which caps local inference as a mass-market proposition. It endures because privacy, offline capability and zero marginal cost are not features an API can offer, and because almost every consumer-facing local AI tool, Ollama included, is built on top of it.

Escape hatches

3 of 3 replacements have a usable free tier.

Agree with the verdict? Cast a vote — or tell us why it is wrong .

more Dev Tools · 19 killed

view all 252 →

also on Meta's list