Tech

I checked whether PostHog should self-host its Haiku calls

I ran 70 of PostHog Desktop's real Haiku prompts through gpt-oss-20b and priced the GPUs.

I’ve been running models on my own machines for a few months, so when Ian Vanagas wrote in a PostHog newsletter about cutting token spend that Desktop’s Haiku calls could one day move to a cheaper, self-hosted model, I immediately wanted to check it myself. Their code is public ofc, so I pulled the exact Haiku prompts out of the Desktop source, ran them on real PostHog issues and commits, and priced what the same calls would cost on rented GPUs.

TLDR for the busy guys:

  • Haiku is a bigger part of the Desktop bill than I initially thought, somewhere between 8% and 36%, and prompt caching cannot bring it down.
  • A hosted gpt-oss-20b wrote task titles and summaries as well as Haiku in a blind comparison, and it is 19x cheaper.
  • Commit messages are where it fell over. It adds bodies the prompt asks it to leave out, and at low effort it slipped Unicode punctuation, mostly non-breaking hyphens, into 29 of its 70 replies.
  • Renting H100s only pays off at about 20 calls a second, around the clock, and PostHog’s own notes already show why that’s hard.

The rest of this post is how I got those numbers. Btw everything is in this gist. Thanks GLM for pushing everything here for me.

what desktop sends to haiku#

Desktop’s LLM gateway client has one line that answers it:

// Bounded helper workloads (titles, summaries, commit messages, PR copy) run on
// the cheapest model rather than the gateway default.
export const HELPER_GATEWAY_MODEL = "claude-haiku-4-5";

Those helpers are the task title and summary, the commit message and PR description, and the agent harness’s web page summaries. When a free-tier token isn’t allowed to use Haiku, the same file sends these helpers to GLM-5.2, an open model, instead. So the swap path already exists in the code.

To get real sizes, I copied the title and commit prompts word for word and ran them on 40 recent PostHog GitHub issues and 30 recent commits from master. The commit prompt cuts each diff at 8,000 characters, the same as Desktop does. My own autocommit-rs, a tiny Rust CLI I use every day for commit messages, does pretty much the same thing with a 10 KB budget, so this part felt very familiar. I sent the Haiku requests through a small logging proxy, so the token counts below are what Anthropic billed.

helper calltokens intokens outcost per call
title and summary1,450106$0.0020
commit message2,88518$0.0030

Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens. Prompt caching can’t help here either. Haiku 4.5 only caches prompts of 4,096 tokens or more, and these calls are smaller than that, so every call pays full price.

how much of the bill this is#

Desktop’s main model is Opus 4.8, and their newsletter gives the ratio:

The Claude Code SDK in PostHog Desktop uses Haiku extensively (3-5 calls per Opus call) revealing an opportunity for us to swap it with a cheaper, self-hosted model in the future.

To price the Opus side, I logged Claude Code 2.1.292 running a small bug fix on Opus 4.8. A typical call read about 40,000 cached tokens and cost $0.027. And with a lot of MCP tools loaded, the context grew to 114,000 tokens and a call cost $0.069. With three to five Haiku calls per Opus call, Haiku comes to 8% to 36% of the Desktop bill, depending on how big the context is.

That’s much more than plain Claude Code. In public usage logs, Haiku was 2.3% and 1.3% of spend. Desktop leans on Haiku far more because of its own helpers, which is probably why it stood out in Ian’s numbers.

so does gpt-oss-20b do the job?#

gpt-oss-20b is the obvious open-weight candidate. It has 21B total parameters, but only 3.6B are used for each token, and the weights take about 13.5 GiB. I ran it on my M4 Pro Max with llama.cpp, using the same 70 prompts, at both low and medium reasoning effort.

First, the rules in PostHog’s own prompts. Each row counts how many outputs followed that rule.

rule from the promptHaiku 4.5gpt-oss-20b, lowgpt-oss-20b, medium
title of six words or fewer15/4034/4040/40
summary written as “The user…“35/4040/4039/40
commit first line of 72 characters or fewer22/306/3026/30
commit message with no extra body29/3015/3018/30
plain ASCII punctuation70/7041/7050/70
same commit type as the human (fix, feat)18/3017/3018/30

Then I asked Opus 4.8 to compare each pair blind, with the two outputs shown in random order. The judge is an Anthropic model, so it may lean toward Haiku, but its reasons matched the rule checks above.

taskHaiku winsgpt-oss-20b medium wins
titles and summaries1921
commit messages264

So titles and summaries are a tie. gpt-oss-20b at medium effort actually follows the title rules better than Haiku does. Here’s one the judge gave to gpt-oss-20b, for an issue about merging two phone headers into one row:

Haiku 4.5:    Fix TodayPhoneHeader layout and merge with scene title
gpt-oss-20b:  Merge phone header and scene title

The prompt asks for six words or fewer, and Haiku used eight. Most of its misses looked like this, a perfectly good title that ran to seven or eight words.

Commit messages went the other way. gpt-oss-20b often adds a bullet list under the subject, even though the prompt says to output only the commit message. This is PostHog commit 033b526b98, which a person wrote as fix(signup): accept a blank role on social signup:

Haiku 4.5:
fix(signup): allow blank role in organization confirmation form

gpt-oss-20b, medium:
feat(signup): make role selection optional on organization confirmation

- Add `showOptional` prop to `SignupRoleSelect` to mark role as optional
- Update `ConfirmOrganization` to pass `showOptional`
- Allow blank `role_at_organization` in social signup serializer
- Add tests for optional role label and for accepting a blank role in the API

Honestly, I’d be happy to see those bullets in a PR description. But the prompt says not to include any explanation, and it turned a bug fix into a feat on the way. It also likes Unicode non-breaking hyphens (‑), which then end up in git history. Haiku produced none in 70 calls. Some of this could be fixed with a stricter prompt and a small cleanup step, but I wouldn’t move commits over until that’s tested.

Medium effort has a cost. gpt-oss-20b thinks before it answers, and at medium it wrote about four times more output tokens than at low: 444 per title call, against 128. Those tokens are billed like any other output.

what the same call costs on a gpu#

To price a GPU, I first needed to know how many calls fit on one at once, and that comes down to the KV cache. The KV cache is the memory a model keeps for every token of every request in progress. For gpt-oss-20b, each layer stores a key and a value for 8 KV heads of 64 numbers each, in 2-byte BF16:

2 (K and V) * 8 KV heads * 64 dims * 2 bytes
  = 2,048 bytes per layer
  * 24 layers
  = 48 KiB per token

Half of those layers only look back 128 tokens, so less memory is in use than this suggests, but vLLM’s capacity numbers match the full 48 KiB. On an A6000 it reported 28.77 GiB of KV cache as 628,512 tokens, which is exactly that. On an H100, a memory profiling run left 50.28 GiB for the KV cache, which is about 1.1 million tokens. That’s room for about 700 title calls at the same time. For comparison, Llama 3.1 8B needs 128 KiB per token, as I worked out in the 64 users post.

So for these calls, compute runs out long before memory does. GPUStack measured gpt-oss-20b on one H100 with vLLM at about 33,000 tokens a second, for both short and long prompts. That’s about 21 title calls a second, or 11 commit calls a second.

An H100 on Modal costs $3.95 an hour. A rented GPU costs the same whether it’s busy or idle, so what matters is how busy it is on average.

title calls per second, all daycost per callcompared to Haiku
21 (the H100 is full)$0.00005238 times cheaper
5$0.000229 times cheaper
0.55$0.0020the same

One H100 running all month on Modal costs about $2,880. To beat Haiku on title calls, it needs about 1.5 million calls a month, every hour of every day. With a second GPU for spare capacity, which I’ll come back to, it’s $5,770 a month and 2.9 million calls. For comparison, my laptop managed 0.39 title calls a second with eight requests at once, so this needs real datacenter GPUs.

a hosted open model costs the same as a full gpu#

Other companies already run gpt-oss-20b on GPUs they keep busy, and sell the tokens. DeepInfra charges $0.03 per million input tokens and $0.14 per million output tokens. With the token counts from my eval:

callHaiku 4.5hosted gpt-oss-20b, mediumcheaper by
title and summary$0.0020$0.0001019 times
commit message$0.0030$0.0001817 times

At low effort the title call costs $0.000059, which is about the same as a fully loaded H100 at $0.000052. So providers are already selling tokens at roughly the price of a GPU that never rests. A GPU PostHog rents for itself only wins if PostHog can keep it fuller than they keep theirs. That’s hard for traffic that follows people’s working hours.

what posthog already learned running vllm#

PostHog already runs open models on its own GPUs. The LLM gateway sends GLM-5.2 and Kimi K3 to vLLM servers on Modal, moves GLM traffic over a fraction at a time, and falls back to Cloudflare. The notes from their reviewer model experiments are public, and they already hit most of what a Haiku swap would run into.

The Modal deployment returned 503 errors when 9 review sessions hit it at the same time on July 24. That’s why I priced a second GPU. A self-hosted model needs spare capacity and a fallback, and the spare capacity sits idle most of the time.

Self-hosted tokens show up as $0.00 in the gateway’s cost tracking, because there’s no per-token price to look up. If the Haiku calls moved to a private GPU, the per-user cost charts would show them as free while the GPU bill grew somewhere else.

Also the caching moved their numbers a lot. In one round, GLM came out 50% to 100% more expensive, and the notes put that down to the Cloudflare route having no prompt caching. The Desktop helper calls are too short to cache on Haiku anyway, but a self-hosted setup would have to get caching right for the main agent calls.

Cold starts matter too. RunPod got a vLLM cold start for a 32B model from 324 seconds down to 91, and Modal got one down to 12 seconds with GPU memory snapshots on a 3B model. A title call takes about two seconds, so the GPU has to stay warm, and a warm GPU bills by the hour.

some conclusions#

If I could I’d move the title and summary helper to a hosted gpt-oss-20b at medium effort, behind the gateway, starting with a small fraction of traffic the same way GLM moved onto Modal. Haiku stays as the fallback. On my numbers that’s 19 times cheaper per call, with output that matched Haiku in a blind comparison. Desktop already falls back to GLM-5.2 for some free-tier helpers, so it would be worth running the same 70 prompts on GLM too before picking one.

I’d keep commit messages and PR descriptions on Haiku for now. Before moving them, I’d tighten the prompt, turn Unicode punctuation into plain ASCII, and cut the subject at 72 characters, then rerun the comparison. PostHog’s own LLM analytics would show how the helper calls split between titles and commits, and that split decides how much of the savings is available today.

I’d only rent GPUs for this once the helper calls reach about 20 a second, around the clock, or if the data can’t leave PostHog’s machines. Until then, paying per token is the cheaper option.

TRIP COMPUTER / SESSION

TRIP A

Current drive.

A private counter for this browser session. Nothing here is transmitted or retained after the session ends.

Elapsed
00:00
Sections
0
Notes
0
Screens
0

Route/

Build plate11bded9
Chassis
v7.1.3
Revision
11bded9
Last serviced
07 Oct 2026

OWNER’S MANUAL / WTHRAJAT

OPERATING NOTES

How this thing moves.

The header behaves like a small mechanical system. Its readings respond to how you move through the site.

Throttle
Scrolling is input. Faster downward movement builds more momentum and engine speed.
Transmission
Upshifts follow sustained input. Scrolling upward slows and downshifts, and may briefly show reverse.
Idle
When input stops, RPM settles near 850 with mechanical drift. The gearbox eventually returns to neutral.
Tachometer
The needle always follows the reported RPM. It is never calculated from your position on the page.
Trip A
Session time, explored sections, opened notes and approximate screens travelled stay in this tab session.

Controls

Ctrl K
Search notes
?
Open this manual
Esc
Close an instrument
Tab
Move through controls