Your AI Agent Retried and Sent the Email Twice. Give Every Action an Idempotency Key.

A cron job times out halfway through a run. The scheduler does the sensible thing and retries it. The agent wakes up with no memory of the first attempt, reads the same task, and sends the same follow-up email to the same customer. Now there are two emails, and one of them says “just checking in” an hour after the first one said it.

Nothing was broken in the usual sense. The model worked. The tool worked. The retry logic worked. The result was still wrong.

Retries make unattended automation survive flaky networks, rate limits, and crashed processes. But a retry is only safe if the action being repeated is safe to repeat. For most real agent actions, it isn’t, unless you make it so.

The agent cannot tell “never ran” from “ran, but I missed the result”

When a call fails, there are three possible worlds:

  1. The request never reached the other side. Retrying is correct.
  2. The request succeeded, but the response got lost. Retrying duplicates the action.
  3. The request partly succeeded. Retrying may duplicate half of it.

From inside the agent, all three look identical: a timeout or an error. Humans resolve this by checking. Agents usually just retry, because the prompt said to be persistent.

The fix is the one payment systems adopted long ago: idempotency. An idempotent action produces the same end state whether you run it once or five times. You don’t make the agent better at guessing which world it is in. You make the question irrelevant.

What an idempotency key is

An idempotency key is a stable identifier for one intended action. Not one attempt. One intent.

  • Attempt ID: run-8841-attempt-2 (changes every retry, useless for dedup)
  • Intent key: followup-email:customer-4417:quote-2026-10-q3 (identical on every retry)

Build the key from the facts that define the action: what kind, on what object, for what reason. If two attempts carry the same key, they are the same action, and the second should be a no-op that returns the first result.

Avoid timestamps down to the second, random UUIDs generated at call time, and the model’s own wording of the task. A UUID generated inside the retry defeats the point. The key has to exist before the first attempt and survive into the second.

The action ledger

You don’t need a distributed system. On a single self-hosted box, a SQLite table covers it. The ledger records intent keys and outcomes:

  • key (unique)
  • status (started, succeeded, failed)
  • started_at, finished_at
  • result_ref (message ID, invoice number, URL: something that proves it happened)

Around every side-effecting action:

  1. Compute the key.
  2. Look it up. If succeeded, skip and return the stored result_ref.
  3. If started and fresh, another run may be in flight. Wait or bail.
  4. If started and stale, or failed, run the verification step below.
  5. Insert started, perform the action, record succeeded with the result reference.

The unique constraint does the heavy lifting. Two concurrent runs cannot both insert the same intent, so the race resolves in the database instead of in a prayer.

The stale started row is where the decisions live

A row stuck at started means the process died somewhere between “about to send” and “recorded that I sent.” Did the email go out? Three honest options, best first:

Ask the destination. Search the sent folder for a message carrying the intent key in a header. List recent invoices for that customer. Check whether the comment is already on the page. This reconciliation is the only approach that truly resolves the ambiguity.

Use the destination’s own idempotency support. Stripe and many payment, ticketing, and email APIs accept an idempotency header. Pass your intent key through and the provider dedupes for you, even across your crashes.

Escalate to a human. If you can’t verify and the action is high-stakes or public, park it in an exception queue with the key, the target, and what you know. A five-second human decision beats an apology email.

What you should not do is default to “assume it failed and retry.” That default is how the double-send happened.

Sort actions by how forgiving they are

Not every action deserves this ceremony.

Naturally idempotent. Setting a field to a value, writing a file to a fixed path, upserting a row by primary key. Run them as often as you like.

Reversible but annoying. Creating a draft, adding a label, posting an internal note. A duplicate is clutter, not damage. A cheap dedup check is enough.

Irreversible or external. Email, SMS, public posts, charges, deletions, webhooks that someone else’s system acts on. These get the full ledger, verification, and a stale-entry policy.

The pain concentrates in that third group, usually a small fraction of what your agent does. Spend the effort there.

Retry rules that cooperate with the ledger

  • Retry the check, not the action. A retry’s first move is consulting the ledger, not calling the tool.
  • Cap attempts and back off. Three attempts with growing delays beats an infinite loop.
  • Separate transient from permanent errors. A 429 or timeout may deserve a retry. A 400 or a revoked token never will.
  • Make “gave up” visible. When attempts run out, write failed and surface it. Silent abandonment is how tasks vanish.
  • Keep the key stable across re-planning. If the agent restarts and rewrites its plan, the key still derives from the facts, not the plan text.

Break it on purpose

Pick one irreversible action in a sandbox and run these drills:

  1. Run it twice back to back. Expect one effect and one skipped duplicate.
  2. Kill the process after the action but before the ledger write, then restart. Expect reconciliation to find the effect and mark it done.
  3. Run two copies at once. Expect exactly one winner.
  4. Simulate a lost response (tool succeeds, agent sees a timeout). Expect no duplicate on retry.

If drill 2 produces a duplicate, you found a real bug before a customer did. That drill alone is worth the afternoon.

The takeaway

Autonomy and retries pull in opposite directions. The more freely an agent recovers from failure, the more chances it has to do the same thing twice. Asking the model to be careful won’t fix that. A boring table with a unique constraint will.

Give every outbound action a key. Write down what happened. When unsure, check the world before repeating yourself. Your agent will look far more competent, and your customers will only ever receive one email.

More from the build log

Suggested

Want the full MarketMai stack?

Get the core MarketMai guides and operator playbooks in one premium bundle for $49.

View Bundle