My bike went for servicing on Saturday and it was not coming back before Monday, so I had a free weekend and nowhere to go. I spent it testing Restate properly instead of just reading its docs.
What Restate promises#
If you haven’t used it, Restate runs your backend code (a handler) and writes
down the result of every step in a journal. You wrap anything that talks to
the outside world, like charging a card or calling an LLM, in ctx.run,
something like this:
const amount = await ctx.run("ask model", () => askModel(order));
await ctx.run("charge card", () => chargeCard(order, amount));
If the process crashes or the connection drops halfway, Restate retries the handler. Steps that already finished are not run again, their saved result is read back from the journal. So the card should be charged once even if things break in the middle.
The docs are honest that there is one exception. If a step finished but its result had not been saved yet, Restate has no way to know it finished, so the retry runs it again. I wanted to see that gap with my own eyes: where exactly it is, how big it is, and what it costs in a real handler.
My first test proved nothing#
On Saturday I built the obvious test. A small agent that bills a card on every turn, running on Restate. I killed the Restate server in the middle of payments, then pushed 250,900 more turns through the same setup and checked that every turn was billed exactly once. It was.
Then I read my own code again. My ledger remembered every key it had already billed, and the Stripe call sent the same key. So even if Restate had run a charge twice, my code or Stripe would have quietly removed the second one. The test could not fail, so it told me nothing about Restate. Killing things at random times has a second problem too. The gap I was looking for is a few milliseconds wide, and a random kill almost never lands inside it.
What I built instead#
Restate talks to your service over one HTTP/2 connection per attempt. Every
message on it is small and simple: an 8 byte header with a type, some flags
and a length, then the body. When a ctx.run starts the SDK sends a Run
message. When the step finishes it sends the result in a
ProposeRunCompletion message, and Restate replies once the result is saved.
Because the messages are this simple, I could put a proxy in the middle that reads them. I called it cutpoint. It works like this:
- Run the handler once with nothing broken, and record every message (12 for my first handler).
- Run the handler again 12 more times. Each time, drop the connection right before a different one of those 12 messages.
- Let Restate retry like it normally would, then count what actually happened outside: how many charges, how many model calls.
The counting part matters most. Restate’s journal will always look fine after a retry, so the only way to catch a double charge is to count it where it happens: at the payments API, the model, or a file on disk.
First test: a checkout#
I started with the simplest checkout I could write. It asks a model how much
to charge, then charges the card. Both steps are in ctx.run and there is no
idempotency key. The model and the payments API are fakes running on my
laptop. They take 400 ms and 250 ms and count every request.
Part of the real output:
$ node src/cli.ts sweep examples/checkout/naive.sweep.ts
baseline: 12 frames, effects {"model calls":1,"model answers":1,"payment requests":1,"cards charged":1}
...
cut 7/12 toRuntime#3 Run (charge card): duplicate {"model calls":1,"model answers":1,"payment requests":2,"cards charged":2} attempts=2
cut 8/12 toRuntime#4 ProposeRunCompletion (charge card): duplicate {"model calls":1,"model answers":1,"payment requests":2,"cards charged":2} attempts=2
cut 9/12 toRuntime#5 AwaitingOn: clean {"model calls":1,"model answers":1,"payment requests":1,"cards charged":1} attempts=2
...
All 12 cuts together:
| where the connection dropped | what happened |
|---|---|
Run of the model call | model answered twice |
ProposeRunCompletion of the model call | model answered twice |
Run of the charge | card charged twice |
ProposeRunCompletion of the charge | card charged twice |
| the other 8 places | same as the normal run |
So there are two places per step, not one.
The ProposeRunCompletion one is the gap from the docs. The card was charged,
the result was on its way, and the connection dropped before Restate saved
it. It is very small. In this run the charge result was saved 3.6 ms after it
was sent. Nothing in your code can close it.
The Run one surprised me. I expected a drop there to be harmless, because
the step had only just started. But a dropped connection does not stop your
code. The old attempt kept going, finished its request to the payments API,
and then had nowhere to send the result. Meanwhile Restate had started a new
attempt, which charged the card again. So this gap is as long as the step
itself: 257 ms for the charge and 410 ms for the model call. That is much
bigger than the 3.6 ms one.
What fixed it#
There are two separate fixes, and they cover different gaps.
The first is an idempotency key on the charge. The payments API then returns the first charge instead of making a new one. That covers both gaps for anything that supports keys.
The second is for the long gap. The TypeScript SDK gives you
ctx.request().attemptCompletedSignal, which fires when an attempt is over.
If you pass it to fetch, the old attempt stops its request instead of
racing the new one. A model call has no idempotency key, and the call itself
is the expensive part, so this matters there.
I wrote the checkout three ways and ran cutpoint on each:
| version | card charged twice at | model answered twice at |
|---|---|---|
| no key, no signal | 2 places | 2 places |
| idempotency key | none | 2 places |
key and attemptCompletedSignal | none | 1 place |
The one left over is the 3.6 ms gap, which no signal can help.
For the model call I was using durableCalls from
@restatedev/vercel-ai-middleware. In 0.4.0 it does not pass the attempt
signal to the model request, so the old call keeps running unless you pass
the signal yourself. I tried a 21 line change that passes it inside the
middleware, and the naive checkout stopped repeating the model answer at the
Run step without any change to the handler. I opened
an issue
for it.
Crashing the service instead#
A dropped connection leaves the old process alive, which is why its code keeps
running. So I also tried killing the service with kill -9 at the same
moment and restarting it. I thought the Run cases would become clean, since
the old attempt dies with the process.
The card was still charged twice. The charge request had already left the process before it died, and the payments API finished it anyway. The model call was different. The model got a second request, but the first one was never answered, so only one answer was generated.
This was the most useful thing I learned. Signals and crashes both stop your own code from doing more work. Neither one can take back a request that already left. Only an idempotency key covers that.
Restate’s own agent#
Then I ran cutpoint on restatedev/agent, Restate’s reference agent, without changing any of its code. It accepts any OpenAI compatible endpoint, so my fake model scripted one turn: run a shell command that adds a line to a file, write another file, then reply. A line in that file means the shell command ran, so counting lines tells me how many times it ran.
One turn is 72 messages, so 72 cuts. I ran the whole thing three times and
every time it found the same one duplicate: the shell command ran twice when
the connection dropped right at its ProposeRunCompletion. That is the same
small gap again. The model also answered twice at its own two
ProposeRunCompletion messages, for the same reason. No Run cut ever
repeated a command, because the agent already passes its signal into the
sandbox and model calls.
I went to the source to see if this was worth an issue, and found this comment right above that tool’s settings:
(Crash recovery can still repeat an unjournaled command.)
So the authors already knew about this exact case. cutpoint found it, and apart from the small gap it found nothing else in 72 cuts. I took that as a good sign that the tool was measuring the right thing.
I also ran a Python version of the checkout through the same proxy, which
works because the proxy only reads messages. It has one extra duplicate. The
Python SDK starts running the step a fraction of a millisecond before it sends
Run, so the long gap starts one message earlier.
What I got out of it#
Before this weekend I knew “a step can run twice” as a sentence from the docs. Now I know where it happens in my own handlers:
- The gap the docs talk about is real and tiny, a few milliseconds.
- The bigger gap is a dropped connection while a step is running. It lasts as
long as the step, and
attemptCompletedSignalcloses most of it. - Neither a signal nor a crash can undo a request that already left, so anything that moves money or sends something still needs an idempotency key.
It also gave me a way to check any handler instead of guessing. A 12 message handler takes about half a minute.
What it does not cover#
cutpoint only drops the connection between Restate and the service. It does not crash the Restate server itself or break its storage, which is what Restate’s own Jepsen tests are for. It drops the connection once per run, so two failures in one run are not covered. The timings are from my laptop and include the proxy’s own hop, so they are good for comparing steps, not for quoting latency. Everything here ran on Restate 1.8.0-rc.1 and the 1.17.2 TypeScript SDK.
The code, all three checkouts, the agent scripts and every report are here: github.com/wthrajat/cutpoint. To test your own handler you write a small file that starts one run and counts what happened outside Restate. The counting is the part that needs real thought.