My bike went for servicing on Saturday and wasn’t coming back before Monday, so I had a free weekend. I had been wanting to try Restate properly, so I spent it testing how it handles failures. The code and all the reports are here: github.com/wthrajat/cutpoint.
If you haven’t used it, Restate runs your backend code and saves the result of every step. If your app crashes or loses its connection halfway, Restate runs the code again, but skips the steps that already finished by reusing their saved results. So every step should happen exactly once.
My goal was to check that. I broke things on purpose in the middle of a request, let Restate recover, and then counted whether anything happened twice or went missing.
What the numbers mean#
- A run is one request to my test app, from start to finish.
- A cut is me dropping the connection between Restate and the app at one moment in that run.
- A wrong result is a run where anything happened twice, anything went missing, the saved state was off, or the reply was different from a normal run with nothing broken.
So “0 wrong results” means Restate recovered correctly every single time.
Results#
Everything ran on my Mac Studio, against one Restate server (1.8.0-rc.1).
First, a test app that only uses things Restate itself is in charge of: saved state, calls between services, messages sent now and later, timers, and waiting for an outside event. If any of these ran twice or went missing, that would be on Restate.
| test | what I broke | runs | wrong results |
|---|---|---|---|
| one cut | the connection, at each of the 52 messages in a run | 49 | 0 |
| two cuts | the same, then again while Restate was recovering | 2,050 | 0 |
| under load | 100 runs at a time, 41,826 random cuts, and the Restate server killed with kill -9 19 times | 20,000 | 0 |
In the load test every request also got its answer, even across the 19 server kills.
Then, a checkout I wrote carelessly on purpose: ask an AI model how much to charge, then charge a card, with no protection. This is where things did run twice, and it was my code each time:
| version | card charged twice | model answered twice |
|---|---|---|
| no protection | at 2 of 12 cuts | at 2 of 12 cuts |
| idempotency key on the charge | never | at 2 of 12 cuts |
| key, and stopping the old run | never | at 1 of 12 cuts |
no protection, app killed with kill -9 | at 2 of 12 cuts | at 1 of 12 cuts |
I also ran restatedev/agent, Restate’s example AI agent, without changing its code. Apart from the model answering again at one tiny gap (more on that below), the only repeat was a shell command, at exactly the moment a comment in its code already warns about.
How I tested it#
Restate and your app talk by sending small messages back and forth. Each message starts with an 8 byte header, so reading them takes only a few lines:
const type = buffer.readUInt16BE(0);
const flags = buffer.readUInt16BE(2);
const bodyLength = buffer.readUInt32BE(4);
So I put a small proxy between Restate and the app. For the first two tests, it runs the code once normally and writes down every message. Then it runs it again once per message, and cuts the connection right before that message. For the load test, it cuts at random instead, while a script kills the Restate server every 20 seconds. After each run, it counts what actually happened outside: how many charges, messages and AI answers.
Why my checkout ran twice#
There are two moments in each step where a cut makes the charge happen again:
The second moment is the one Restate’s docs talk about. The card was charged, but the connection dropped before Restate saved the result. It’s tiny, the result was saved 3.6 ms after it was sent.
The first moment surprised me more. Cutting the connection doesn’t stop your code. The old run kept going and finished the charge, while Restate had already started a new run that charged again. So this gap is as long as the charge itself, 257 ms here.
An idempotency key on the charge (an ID that tells the payment API “this is
the same charge as before”) fixes both. Restate’s TypeScript SDK also has
ctx.request().attemptCompletedSignal, which fires when a run is over. If you
pass it to fetch, the old run stops instead of racing the new one. That
matters most for AI calls, since those have no idempotency key. Killing the
app with kill -9 didn’t save the card, because the charge request had
already left the app.
The docs don’t mention attemptCompletedSignal yet, and durableCalls in
Restate’s Vercel AI middleware (0.4.0) doesn’t pass it to the model call, so
there you have to pass it yourself through abortSignal.
One small issue#
On 1.8.0-rc.1, a retry that should wait 2 seconds or more can start up to about 2 seconds early. With the default settings, the third retry should wait at least 2 seconds, and it was starting after about 1.1. It comes down to one line that rounds the time down instead of up. It only changes timing. Nothing ran twice or went missing because of it.
What this doesn’t cover#
This was about correctness while things break, not speed. The load test ran about 48 runs a second on my Mac Studio, which is far less than what real Restate users run. It was one server, not a cluster, so it doesn’t cover failover or lost disks. Restate’s own Jepsen tests cover those. The AI model and the payment API were local fakes with fixed delays.
Wrapping up#
I went in trying to break it and mostly ended up learning how it works. In about 22,000 runs, nothing I did broke Restate’s side of things. The things that did run twice were in my own code, and Restate already has a fix for both.
The proxy, the test apps and every report are here: github.com/wthrajat/cutpoint.