Tech

I broke Restate's connection at every step to see what runs twice

A small weekend tool that drops the connection between Restate and my service at every possible point, then counts which charges and model calls happened twice.

My bike went for servicing on Saturday and it was not coming back before Monday, so I had a free weekend and nowhere to go. I spent it testing Restate properly instead of just reading its docs.

What Restate promises#

If you haven’t used it, Restate runs your backend code (a handler) and writes down the result of every step in a journal. You wrap anything that talks to the outside world, like charging a card or calling an LLM, in ctx.run, something like this:

const amount = await ctx.run("ask model", () => askModel(order));
await ctx.run("charge card", () => chargeCard(order, amount));

If the process crashes or the connection drops halfway, Restate retries the handler. Steps that already finished are not run again, their saved result is read back from the journal. So the card should be charged once even if things break in the middle.

The docs are honest that there is one exception. If a step finished but its result had not been saved yet, Restate has no way to know it finished, so the retry runs it again. I wanted to see that gap with my own eyes: where exactly it is, how big it is, and what it costs in a real handler.

My first test proved nothing#

On Saturday I built the obvious test. A small agent that bills a card on every turn, running on Restate. I killed the Restate server in the middle of payments, then pushed 250,900 more turns through the same setup and checked that every turn was billed exactly once. It was.

Then I read my own code again. My ledger remembered every key it had already billed, and the Stripe call sent the same key. So even if Restate had run a charge twice, my code or Stripe would have quietly removed the second one. The test could not fail, so it told me nothing about Restate. Killing things at random times has a second problem too. The gap I was looking for is a few milliseconds wide, and a random kill almost never lands inside it.

What I built instead#

Restate talks to your service over one HTTP/2 connection per attempt. Every message on it is small and simple: an 8 byte header with a type, some flags and a length, then the body. When a ctx.run starts the SDK sends a Run message. When the step finishes it sends the result in a ProposeRunCompletion message, and Restate replies once the result is saved.

Because the messages are this simple, I could put a proxy in the middle that reads them. I called it cutpoint. It works like this:

  1. Run the handler once with nothing broken, and record every message (12 for my first handler).
  2. Run the handler again 12 more times. Each time, drop the connection right before a different one of those 12 messages.
  3. Let Restate retry like it normally would, then count what actually happened outside: how many charges, how many model calls.
counts effects restate-server cutpoint your service outside world

The counting part matters most. Restate’s journal will always look fine after a retry, so the only way to catch a double charge is to count it where it happens: at the payments API, the model, or a file on disk.

First test: a checkout#

I started with the simplest checkout I could write. It asks a model how much to charge, then charges the card. Both steps are in ctx.run and there is no idempotency key. The model and the payments API are fakes running on my laptop. They take 400 ms and 250 ms and count every request.

Part of the real output:

$ node src/cli.ts sweep examples/checkout/naive.sweep.ts
baseline: 12 frames, effects {"model calls":1,"model answers":1,"payment requests":1,"cards charged":1}
...
cut 7/12 toRuntime#3 Run (charge card): duplicate {"model calls":1,"model answers":1,"payment requests":2,"cards charged":2} attempts=2
cut 8/12 toRuntime#4 ProposeRunCompletion (charge card): duplicate {"model calls":1,"model answers":1,"payment requests":2,"cards charged":2} attempts=2
cut 9/12 toRuntime#5 AwaitingOn: clean {"model calls":1,"model answers":1,"payment requests":1,"cards charged":1} attempts=2
...

All 12 cuts together:

where the connection droppedwhat happened
Run of the model callmodel answered twice
ProposeRunCompletion of the model callmodel answered twice
Run of the chargecard charged twice
ProposeRunCompletion of the chargecard charged twice
the other 8 placessame as the normal run

So there are two places per step, not one.

Run "charge card" charge charged ProposeRunCompletion Ack, result stored cut here and the old attempt keeps going cut here and the result is lost restate-server your service payments API

The ProposeRunCompletion one is the gap from the docs. The card was charged, the result was on its way, and the connection dropped before Restate saved it. It is very small. In this run the charge result was saved 3.6 ms after it was sent. Nothing in your code can close it.

The Run one surprised me. I expected a drop there to be harmless, because the step had only just started. But a dropped connection does not stop your code. The old attempt kept going, finished its request to the payments API, and then had nowhere to send the result. Meanwhile Restate had started a new attempt, which charged the card again. So this gap is as long as the step itself: 257 ms for the charge and 410 ms for the model call. That is much bigger than the 3.6 ms one.

What fixed it#

There are two separate fixes, and they cover different gaps.

The first is an idempotency key on the charge. The payments API then returns the first charge instead of making a new one. That covers both gaps for anything that supports keys.

The second is for the long gap. The TypeScript SDK gives you ctx.request().attemptCompletedSignal, which fires when an attempt is over. If you pass it to fetch, the old attempt stops its request instead of racing the new one. A model call has no idempotency key, and the call itself is the expensive part, so this matters there.

I wrote the checkout three ways and ran cutpoint on each:

versioncard charged twice atmodel answered twice at
no key, no signal2 places2 places
idempotency keynone2 places
key and attemptCompletedSignalnone1 place

The one left over is the 3.6 ms gap, which no signal can help.

For the model call I was using durableCalls from @restatedev/vercel-ai-middleware. In 0.4.0 it does not pass the attempt signal to the model request, so the old call keeps running unless you pass the signal yourself. I tried a 21 line change that passes it inside the middleware, and the naive checkout stopped repeating the model answer at the Run step without any change to the handler. I opened an issue for it.

Crashing the service instead#

A dropped connection leaves the old process alive, which is why its code keeps running. So I also tried killing the service with kill -9 at the same moment and restarting it. I thought the Run cases would become clean, since the old attempt dies with the process.

The card was still charged twice. The charge request had already left the process before it died, and the payments API finished it anyway. The model call was different. The model got a second request, but the first one was never answered, so only one answer was generated.

This was the most useful thing I learned. Signals and crashes both stop your own code from doing more work. Neither one can take back a request that already left. Only an idempotency key covers that.

Restate’s own agent#

Then I ran cutpoint on restatedev/agent, Restate’s reference agent, without changing any of its code. It accepts any OpenAI compatible endpoint, so my fake model scripted one turn: run a shell command that adds a line to a file, write another file, then reply. A line in that file means the shell command ran, so counting lines tells me how many times it ran.

One turn is 72 messages, so 72 cuts. I ran the whole thing three times and every time it found the same one duplicate: the shell command ran twice when the connection dropped right at its ProposeRunCompletion. That is the same small gap again. The model also answered twice at its own two ProposeRunCompletion messages, for the same reason. No Run cut ever repeated a command, because the agent already passes its signal into the sandbox and model calls.

I went to the source to see if this was worth an issue, and found this comment right above that tool’s settings:

(Crash recovery can still repeat an unjournaled command.)

So the authors already knew about this exact case. cutpoint found it, and apart from the small gap it found nothing else in 72 cuts. I took that as a good sign that the tool was measuring the right thing.

I also ran a Python version of the checkout through the same proxy, which works because the proxy only reads messages. It has one extra duplicate. The Python SDK starts running the step a fraction of a millisecond before it sends Run, so the long gap starts one message earlier.

What I got out of it#

Before this weekend I knew “a step can run twice” as a sentence from the docs. Now I know where it happens in my own handlers:

  • The gap the docs talk about is real and tiny, a few milliseconds.
  • The bigger gap is a dropped connection while a step is running. It lasts as long as the step, and attemptCompletedSignal closes most of it.
  • Neither a signal nor a crash can undo a request that already left, so anything that moves money or sends something still needs an idempotency key.

It also gave me a way to check any handler instead of guessing. A 12 message handler takes about half a minute.

What it does not cover#

cutpoint only drops the connection between Restate and the service. It does not crash the Restate server itself or break its storage, which is what Restate’s own Jepsen tests are for. It drops the connection once per run, so two failures in one run are not covered. The timings are from my laptop and include the proxy’s own hop, so they are good for comparing steps, not for quoting latency. Everything here ran on Restate 1.8.0-rc.1 and the 1.17.2 TypeScript SDK.

The code, all three checkouts, the agent scripts and every report are here: github.com/wthrajat/cutpoint. To test your own handler you write a small file that starts one run and counts what happened outside Restate. The counting is the part that needs real thought.

TRIP COMPUTER / SESSION

TRIP A

Current drive.

A private counter for this browser session. Nothing here is transmitted or retained after the session ends.

Elapsed
00:00
Sections
0
Notes
0
Screens
0

Route/

Build plateb898ff9
Chassis
v7.1.3
Revision
b898ff9
Last serviced
04 Oct 2026

OWNER’S MANUAL / WTHRAJAT

OPERATING NOTES

How this thing moves.

The header behaves like a small mechanical system. Its readings respond to how you move through the site.

Throttle
Scrolling is input. Faster downward movement builds more momentum and engine speed.
Transmission
Upshifts follow sustained input. Scrolling upward slows and downshifts, and may briefly show reverse.
Idle
When input stops, RPM settles near 850 with mechanical drift. The gearbox eventually returns to neutral.
Tachometer
The needle always follows the reported RPM. It is never calculated from your position on the page.
Trip A
Session time, explored sections, opened notes and approximate screens travelled stay in this tab session.

Controls

Ctrl K
Search notes
?
Open this manual
Esc
Close an instrument
Tab
Move through controls