Tech

Testing how Restate handles failures

Cutting connections and killing the server, then counting what happened.

My bike went for servicing on Saturday and wasn’t coming back before Monday, so I had a free weekend. I had been wanting to try Restate properly, so I spent it testing how it handles failures. The code and all the reports are here: github.com/wthrajat/cutpoint.

If you haven’t used it, Restate runs your backend code and saves the result of every step. If your app crashes or loses its connection halfway, Restate runs the code again, but skips the steps that already finished by reusing their saved results. So every step should happen exactly once.

My goal was to check that. I broke things on purpose in the middle of a request, let Restate recover, and then counted whether anything happened twice or went missing.

What the numbers mean#

  • A run is one request to my test app, from start to finish.
  • A cut is me dropping the connection between Restate and the app at one moment in that run.
  • A wrong result is a run where anything happened twice, anything went missing, the saved state was off, or the reply was different from a normal run with nothing broken.

So “0 wrong results” means Restate recovered correctly every single time.

Results#

Everything ran on my Mac Studio, against one Restate server (1.8.0-rc.1).

First, a test app that only uses things Restate itself is in charge of: saved state, calls between services, messages sent now and later, timers, and waiting for an outside event. If any of these ran twice or went missing, that would be on Restate.

testwhat I brokerunswrong results
one cutthe connection, at each of the 52 messages in a run490
two cutsthe same, then again while Restate was recovering2,0500
under load100 runs at a time, 41,826 random cuts, and the Restate server killed with kill -9 19 times20,0000

In the load test every request also got its answer, even across the 19 server kills.

Then, a checkout I wrote carelessly on purpose: ask an AI model how much to charge, then charge a card, with no protection. This is where things did run twice, and it was my code each time:

versioncard charged twicemodel answered twice
no protectionat 2 of 12 cutsat 2 of 12 cuts
idempotency key on the chargeneverat 2 of 12 cuts
key, and stopping the old runneverat 1 of 12 cuts
no protection, app killed with kill -9at 2 of 12 cutsat 1 of 12 cuts

I also ran restatedev/agent, Restate’s example AI agent, without changing its code. Apart from the model answering again at one tiny gap (more on that below), the only repeat was a shell command, at exactly the moment a comment in its code already warns about.

How I tested it#

Restate and your app talk by sending small messages back and forth. Each message starts with an 8 byte header, so reading them takes only a few lines:

const type = buffer.readUInt16BE(0);
const flags = buffer.readUInt16BE(2);
const bodyLength = buffer.readUInt32BE(4);

So I put a small proxy between Restate and the app. For the first two tests, it runs the code once normally and writes down every message. Then it runs it again once per message, and cuts the connection right before that message. For the load test, it cuts at random instead, while a script kills the Restate server every 20 seconds. After each run, it counts what actually happened outside: how many charges, messages and AI answers.

counts Restate proxy my app outside world

Why my checkout ran twice#

There are two moments in each step where a cut makes the charge happen again:

starting "charge card" charge charged here is the result saved cut here and the old run keeps going cut here and the result is lost Restate my app payment API

The second moment is the one Restate’s docs talk about. The card was charged, but the connection dropped before Restate saved the result. It’s tiny, the result was saved 3.6 ms after it was sent.

The first moment surprised me more. Cutting the connection doesn’t stop your code. The old run kept going and finished the charge, while Restate had already started a new run that charged again. So this gap is as long as the charge itself, 257 ms here.

An idempotency key on the charge (an ID that tells the payment API “this is the same charge as before”) fixes both. Restate’s TypeScript SDK also has ctx.request().attemptCompletedSignal, which fires when a run is over. If you pass it to fetch, the old run stops instead of racing the new one. That matters most for AI calls, since those have no idempotency key. Killing the app with kill -9 didn’t save the card, because the charge request had already left the app.

The docs don’t mention attemptCompletedSignal yet, and durableCalls in Restate’s Vercel AI middleware (0.4.0) doesn’t pass it to the model call, so there you have to pass it yourself through abortSignal.

One small issue#

On 1.8.0-rc.1, a retry that should wait 2 seconds or more can start up to about 2 seconds early. With the default settings, the third retry should wait at least 2 seconds, and it was starting after about 1.1. It comes down to one line that rounds the time down instead of up. It only changes timing. Nothing ran twice or went missing because of it.

What this doesn’t cover#

This was about correctness while things break, not speed. The load test ran about 48 runs a second on my Mac Studio, which is far less than what real Restate users run. It was one server, not a cluster, so it doesn’t cover failover or lost disks. Restate’s own Jepsen tests cover those. The AI model and the payment API were local fakes with fixed delays.

Wrapping up#

I went in trying to break it and mostly ended up learning how it works. In about 22,000 runs, nothing I did broke Restate’s side of things. The things that did run twice were in my own code, and Restate already has a fix for both.

The proxy, the test apps and every report are here: github.com/wthrajat/cutpoint.

TRIP COMPUTER / SESSION

TRIP A

Current drive.

A private counter for this browser session. Nothing here is transmitted or retained after the session ends.

Elapsed
00:00
Sections
0
Notes
0
Screens
0

Route/

Build plate74b649a
Chassis
v7.1.3
Revision
74b649a
Last serviced
06 Oct 2026

OWNER’S MANUAL / WTHRAJAT

OPERATING NOTES

How this thing moves.

The header behaves like a small mechanical system. Its readings respond to how you move through the site.

Throttle
Scrolling is input. Faster downward movement builds more momentum and engine speed.
Transmission
Upshifts follow sustained input. Scrolling upward slows and downshifts, and may briefly show reverse.
Idle
When input stops, RPM settles near 850 with mechanical drift. The gearbox eventually returns to neutral.
Tachometer
The needle always follows the reported RPM. It is never calculated from your position on the page.
Trip A
Session time, explored sections, opened notes and approximate screens travelled stay in this tab session.

Controls

Ctrl K
Search notes
?
Open this manual
Esc
Close an instrument
Tab
Move through controls