I’ve been running models on my own machines for a few months, so when Ian Vanagas wrote in a PostHog newsletter about cutting token spend that Desktop’s Haiku calls could one day move to a cheaper, self-hosted model, I immediately wanted to check it myself. Their code is public ofc, so I pulled the exact Haiku prompts out of the Desktop source, ran them on real PostHog issues and commits, and priced what the same calls would cost on rented GPUs.
TLDR for the busy guys:
- Haiku is a bigger part of the Desktop bill than I initially thought, somewhere between 8% and 36%, and prompt caching cannot bring it down.
- A hosted
gpt-oss-20bwrote task titles and summaries as well as Haiku in a blind comparison, and it is 19x cheaper. - Commit messages are where it fell over. It adds bodies the prompt asks it to leave out, and at low effort it slipped Unicode punctuation, mostly non-breaking hyphens, into 29 of its 70 replies.
- Renting H100s only pays off at about 20 calls a second, around the clock, and PostHog’s own notes already show why that’s hard.
The rest of this post is how I got those numbers. Btw everything is in this gist. Thanks GLM for pushing everything here for me.
what desktop sends to haiku#
Desktop’s LLM gateway client has one line that answers it:
// Bounded helper workloads (titles, summaries, commit messages, PR copy) run on
// the cheapest model rather than the gateway default.
export const HELPER_GATEWAY_MODEL = "claude-haiku-4-5";
Those helpers are the task title and summary, the commit message and PR description, and the agent harness’s web page summaries. When a free-tier token isn’t allowed to use Haiku, the same file sends these helpers to GLM-5.2, an open model, instead. So the swap path already exists in the code.
To get real sizes, I copied the title and commit prompts word for word and ran
them on 40 recent PostHog GitHub issues and 30 recent commits from master.
The commit prompt cuts each diff at 8,000 characters, the same as Desktop does.
My own autocommit-rs, a tiny Rust
CLI I use every day for commit messages, does pretty much the same thing with a
10 KB budget, so this part felt very familiar.
I sent the Haiku requests through a small logging proxy, so the token counts
below are what Anthropic billed.
| helper call | tokens in | tokens out | cost per call |
|---|---|---|---|
| title and summary | 1,450 | 106 | $0.0020 |
| commit message | 2,885 | 18 | $0.0030 |
Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens. Prompt caching can’t help here either. Haiku 4.5 only caches prompts of 4,096 tokens or more, and these calls are smaller than that, so every call pays full price.
how much of the bill this is#
Desktop’s main model is Opus 4.8, and their newsletter gives the ratio:
The Claude Code SDK in PostHog Desktop uses Haiku extensively (3-5 calls per Opus call) revealing an opportunity for us to swap it with a cheaper, self-hosted model in the future.
To price the Opus side, I logged Claude Code 2.1.292 running a small bug fix on Opus 4.8. A typical call read about 40,000 cached tokens and cost $0.027. And with a lot of MCP tools loaded, the context grew to 114,000 tokens and a call cost $0.069. With three to five Haiku calls per Opus call, Haiku comes to 8% to 36% of the Desktop bill, depending on how big the context is.
That’s much more than plain Claude Code. In public usage logs, Haiku was 2.3% and 1.3% of spend. Desktop leans on Haiku far more because of its own helpers, which is probably why it stood out in Ian’s numbers.
so does gpt-oss-20b do the job?#
gpt-oss-20b is the obvious
open-weight candidate. It has 21B total parameters, but only 3.6B are used for
each token, and the weights take about 13.5 GiB. I ran it on my M4 Pro Max
with llama.cpp, using the same 70 prompts, at both low and medium reasoning
effort.
First, the rules in PostHog’s own prompts. Each row counts how many outputs followed that rule.
| rule from the prompt | Haiku 4.5 | gpt-oss-20b, low | gpt-oss-20b, medium |
|---|---|---|---|
| title of six words or fewer | 15/40 | 34/40 | 40/40 |
| summary written as “The user…“ | 35/40 | 40/40 | 39/40 |
| commit first line of 72 characters or fewer | 22/30 | 6/30 | 26/30 |
| commit message with no extra body | 29/30 | 15/30 | 18/30 |
| plain ASCII punctuation | 70/70 | 41/70 | 50/70 |
same commit type as the human (fix, feat) | 18/30 | 17/30 | 18/30 |
Then I asked Opus 4.8 to compare each pair blind, with the two outputs shown in random order. The judge is an Anthropic model, so it may lean toward Haiku, but its reasons matched the rule checks above.
| task | Haiku wins | gpt-oss-20b medium wins |
|---|---|---|
| titles and summaries | 19 | 21 |
| commit messages | 26 | 4 |
So titles and summaries are a tie. gpt-oss-20b at medium effort actually
follows the title rules better than Haiku does. Here’s one the judge gave to
gpt-oss-20b, for an issue about merging two phone headers into one row:
Haiku 4.5: Fix TodayPhoneHeader layout and merge with scene title
gpt-oss-20b: Merge phone header and scene title
The prompt asks for six words or fewer, and Haiku used eight. Most of its misses looked like this, a perfectly good title that ran to seven or eight words.
Commit messages went the other way. gpt-oss-20b often adds a bullet list
under the subject, even though the prompt says to output only the commit
message. This is PostHog commit 033b526b98, which a person wrote as
fix(signup): accept a blank role on social signup:
Haiku 4.5:
fix(signup): allow blank role in organization confirmation form
gpt-oss-20b, medium:
feat(signup): make role selection optional on organization confirmation
- Add `showOptional` prop to `SignupRoleSelect` to mark role as optional
- Update `ConfirmOrganization` to pass `showOptional`
- Allow blank `role_at_organization` in social signup serializer
- Add tests for optional role label and for accepting a blank role in the API
Honestly, I’d be happy to see those bullets in a PR description. But the prompt
says not to include any explanation, and it turned a bug fix into a feat on
the way. It also likes Unicode non-breaking hyphens (‑), which then end up in
git history. Haiku produced none in 70 calls. Some of this could be fixed with a
stricter prompt and a small cleanup step, but I wouldn’t move commits over until
that’s tested.
Medium effort has a cost. gpt-oss-20b thinks before it answers, and at medium
it wrote about four times more output tokens than at low: 444 per title call,
against 128. Those tokens are billed like any other output.
what the same call costs on a gpu#
To price a GPU, I first needed to know how many calls fit on one at once, and
that comes down to the KV cache. The KV cache is the memory a model keeps for
every token of every request in progress. For gpt-oss-20b, each layer
stores a key and a value for 8 KV heads of 64 numbers each, in 2-byte BF16:
2 (K and V) * 8 KV heads * 64 dims * 2 bytes
= 2,048 bytes per layer
* 24 layers
= 48 KiB per token
Half of those layers only look back 128 tokens, so less memory is in use than this suggests, but vLLM’s capacity numbers match the full 48 KiB. On an A6000 it reported 28.77 GiB of KV cache as 628,512 tokens, which is exactly that. On an H100, a memory profiling run left 50.28 GiB for the KV cache, which is about 1.1 million tokens. That’s room for about 700 title calls at the same time. For comparison, Llama 3.1 8B needs 128 KiB per token, as I worked out in the 64 users post.
So for these calls, compute runs out long before memory does.
GPUStack measured
gpt-oss-20b on one H100 with vLLM at about 33,000 tokens a second, for both
short and long prompts. That’s about 21 title calls a second, or 11 commit
calls a second.
An H100 on Modal costs $3.95 an hour. A rented GPU costs the same whether it’s busy or idle, so what matters is how busy it is on average.
| title calls per second, all day | cost per call | compared to Haiku |
|---|---|---|
| 21 (the H100 is full) | $0.000052 | 38 times cheaper |
| 5 | $0.00022 | 9 times cheaper |
| 0.55 | $0.0020 | the same |
One H100 running all month on Modal costs about $2,880. To beat Haiku on title calls, it needs about 1.5 million calls a month, every hour of every day. With a second GPU for spare capacity, which I’ll come back to, it’s $5,770 a month and 2.9 million calls. For comparison, my laptop managed 0.39 title calls a second with eight requests at once, so this needs real datacenter GPUs.
a hosted open model costs the same as a full gpu#
Other companies already run gpt-oss-20b on GPUs they keep busy, and sell the
tokens. DeepInfra charges $0.03 per
million input tokens and $0.14 per million output tokens. With the token counts
from my eval:
| call | Haiku 4.5 | hosted gpt-oss-20b, medium | cheaper by |
|---|---|---|---|
| title and summary | $0.0020 | $0.00010 | 19 times |
| commit message | $0.0030 | $0.00018 | 17 times |
At low effort the title call costs $0.000059, which is about the same as a fully loaded H100 at $0.000052. So providers are already selling tokens at roughly the price of a GPU that never rests. A GPU PostHog rents for itself only wins if PostHog can keep it fuller than they keep theirs. That’s hard for traffic that follows people’s working hours.
what posthog already learned running vllm#
PostHog already runs open models on its own GPUs. The LLM gateway sends GLM-5.2 and Kimi K3 to vLLM servers on Modal, moves GLM traffic over a fraction at a time, and falls back to Cloudflare. The notes from their reviewer model experiments are public, and they already hit most of what a Haiku swap would run into.
The Modal deployment returned 503 errors when 9 review sessions hit it at the same time on July 24. That’s why I priced a second GPU. A self-hosted model needs spare capacity and a fallback, and the spare capacity sits idle most of the time.
Self-hosted tokens show up as $0.00 in the gateway’s cost tracking, because there’s no per-token price to look up. If the Haiku calls moved to a private GPU, the per-user cost charts would show them as free while the GPU bill grew somewhere else.
Also the caching moved their numbers a lot. In one round, GLM came out 50% to 100% more expensive, and the notes put that down to the Cloudflare route having no prompt caching. The Desktop helper calls are too short to cache on Haiku anyway, but a self-hosted setup would have to get caching right for the main agent calls.
Cold starts matter too. RunPod got a vLLM cold start for a 32B model from 324 seconds down to 91, and Modal got one down to 12 seconds with GPU memory snapshots on a 3B model. A title call takes about two seconds, so the GPU has to stay warm, and a warm GPU bills by the hour.
some conclusions#
If I could I’d move the title and summary helper to a hosted gpt-oss-20b at medium effort,
behind the gateway, starting with a small fraction of traffic the same way GLM
moved onto Modal. Haiku stays as the fallback. On my numbers that’s 19 times
cheaper per call, with output that matched Haiku in a blind comparison. Desktop
already falls back to GLM-5.2 for some free-tier helpers, so it would be worth
running the same 70 prompts on GLM too before picking one.
I’d keep commit messages and PR descriptions on Haiku for now. Before moving them, I’d tighten the prompt, turn Unicode punctuation into plain ASCII, and cut the subject at 72 characters, then rerun the comparison. PostHog’s own LLM analytics would show how the helper calls split between titles and commits, and that split decides how much of the savings is available today.
I’d only rent GPUs for this once the helper calls reach about 20 a second, around the clock, or if the data can’t leave PostHog’s machines. Until then, paying per token is the cheaper option.