The request timed out on the client, but it had already worked on the server. That one sentence covers most of the duplicate-write bugs we find in audits. A phone on a weak signal sends POST /orders. The server takes the payment, writes the order row and starts sending the 201. The connection dies before the response reaches the phone. The Dio client waits out its receive timeout, the retry interceptor does what it was built to do and sends the request again, and the server charges the user a second time.

Neither side is buggy on its own. The retry logic is reasonable and so is the endpoint. The bug is that the two sides never agreed on what a retry means. We now write that agreement down as a contract and require it before any payment or inventory endpoint ships.

How the duplicate actually happens

The client code usually looks harmless. Here’s a typical interceptor:

class RetryInterceptor extends Interceptor {
  @override
  Future<void> onError(DioException err, ErrorInterceptorHandler handler) async {
    final attempt = (err.requestOptions.extra['attempt'] as int?) ?? 0;
    final retryable = err.type == DioExceptionType.receiveTimeout ||
        err.type == DioExceptionType.connectionError;
    if (retryable && attempt < 3) {
      err.requestOptions.extra['attempt'] = attempt + 1;
      return handler.resolve(await dio.fetch(err.requestOptions));
    }
    handler.next(err);
  }
}

The problem is receiveTimeout. A connect timeout means the request almost certainly never arrived. A receive timeout means the request was sent and you don’t know what the server did with it. Retrying a GET after a receive timeout is fine. Retrying a POST that way is a guess, and on a payment endpoint a wrong guess costs the user real money.

The interceptor isn’t the only way to get a duplicate. We also see:

  • Double taps. The pay button stays enabled while the first request is in flight because the loading state only updates after the await.
  • Resubmitting after the app restarts. The OS kills the app mid-request, and on relaunch an offline queue replays the pending action, which may already have gone through.
  • Load balancer or proxy retries. Some infrastructure resends upstream on its own when a backend is slow, and the mobile code has no say in it.

Disabling the interceptor for POSTs fixes one of these four cases at most. The only fix that covers all of them is on the server.

The client half: one key per intent, not per attempt

The client generates an idempotency key when the user decides to do something, not when a request goes out. This is where teams usually get it wrong, even after they’ve decided to use keys. If the key is generated in the interceptor or the API wrapper, each retry gets a new key, and the server sees every retry as a new operation.

// Created when the user confirms checkout, stored with the pending action
final action = PendingCheckout(
  cartId: cart.id,
  idempotencyKey: const Uuid().v4(),
);
await pendingStore.save(action); // survives process death

await dio.post(
  '/v1/orders',
  data: action.toJson(),
  options: Options(headers: {'Idempotency-Key': action.idempotencyKey}),
);

The rules on the client:

  • The key is a random UUID, created once for each user intent and saved before the first attempt.
  • Every retry sends the same key, whether it comes from the interceptor, a double tap or an offline-queue replay.
  • The key is deleted only after a final response: success, or an error the client isn’t allowed to retry.
  • If the user edits the cart, that’s a new intent and gets a new key.

The server half: the dedupe table

On the server we use a single table, scoped to the user so one caller can never replay another caller’s key:

CREATE TABLE idempotency_keys (
  user_id        bigint      NOT NULL,
  key            uuid        NOT NULL,
  endpoint       text        NOT NULL,
  request_hash   bytea       NOT NULL,
  status         text        NOT NULL CHECK (status IN ('in_progress', 'completed')),
  response_code  int,
  response_body  jsonb,
  locked_until   timestamptz,
  created_at     timestamptz NOT NULL DEFAULT now(),
  PRIMARY KEY (user_id, key)
);

The handler does this, in this order:

  1. Claim the key. Run INSERT ... ON CONFLICT DO NOTHING RETURNING with status in_progress. The primary key makes this atomic, so two concurrent requests can’t both claim it.
  2. If the insert did nothing, read the existing row. If request_hash doesn’t match, return 422, because the client reused a key for a different payload, which is a client bug. If the row is in_progress and the lock hasn’t expired, return 409 with Retry-After. If it’s completed, return the stored status code and body exactly as they were first sent.
  3. Do the work. Write the order and mark the key completed with the response in the same database transaction. Then no code path can write the order without also recording the key.
  4. Pass the key to the payment provider. The external charge isn’t part of your transaction. Most payment providers accept their own idempotency key, so send one derived from yours, for example orders:{user_id}:{key}. If your process crashes after the charge but before the commit, the retry reaches the provider with the same key, and the provider returns the original charge instead of creating a new one.

locked_until handles the unpleasant case where a worker claims the key and then dies. When the lock expires, a retry can take over the row. That only works because step 4 already makes the external side effect safe to repeat.

Things we deliberately don’t do

  • We don’t dedupe by comparing payloads alone. Two identical orders a minute apart may both be real.
  • We don’t keep the table in a cache that can evict entries. If the dedupe record disappears while the client is still retrying, the protection is gone. It lives in the primary database, and a cleanup job deletes rows older than the longest time a client could still be retrying, plus a margin.
  • We don’t cache 5xx responses as completed. If the operation rolled back, a retry should be allowed to try again.

The contract, written down

Every endpoint that moves money or stock now ships with a short note in its API doc that answers these questions:

  • Is the endpoint idempotent by itself (PUT on a known ID), or does it require Idempotency-Key?
  • Which responses can the client retry, and does it keep the same key when it does?
  • How long does the server keep keys, and is that longer than the client’s retry window?
  • Which key is passed to each external side effect?

If a PR can’t answer those questions, it doesn’t merge. None of this is new; payment APIs have documented the pattern for years. What goes wrong is the gap between the mobile team and the backend team, where each assumes the other handles retries.

We build Flutter apps and the backends behind them together, so this is one of the first things we check, both in our own work and when we audit someone else’s code. If your retry logic and your POST endpoints were built without anyone agreeing on how retries should work, it’s worth looking at them together before a user on a bad connection finds the gap.