All posts Use cases Services Contact
Login Get started

Making an agent reliable once it runs unattended

It worked in the demo and it fails at 2am on a Sunday. Almost none of that is the model. Here is the list we keep on hand.

Once an agent stops being a demo and starts being infrastructure, the failure modes change character. They stop being about model quality and start being about state, retries and ownership. That is good news: these are boring, structural and fixable, and they are the same handful every time.

1. State that leaks between runs

An agent that remembers something between conversations needs somewhere to put it. The default outcome is a JSON blob only one person understands. Within a month you cannot answer why it did that, and a prompt change silently rewrites history.

Give every run a session id, store everything under it, treat the store as append-only. When a run goes wrong you read the whole trace instead of guessing from a snapshot. It costs a few more reads and buys every incident investigation you have ever had to do.

2. Retries that are not idempotent

This is the one that costs real money. A tool call charges a card, sends an email, writes to a contract. The network times out. The agent retries. You now have two charges.

Every non-read tool call gets an idempotency key derived from the run id and the call index. Anything that moves value defaults to no-retry, surface-it. Read the rest of this in what actually breaks when an agent runs unattended.

3. Nobody owns the alert

Plenty of agents have a dashboard. Very few have an alert with a name attached. A dashboard nobody watches is a museum — the point of monitoring is not the chart, it is the person who is accountable when it fires.

So build the console around the alert, not around the telemetry. What broke, when, which run, what it was about to do, and one button to stop it. Monitoring from $699; if you already have the traces and need the surface, that is the console work.

4. Nobody can interrupt it

An agent that cannot be stopped is not autonomous, it is unsupervised. Human-in-the-loop workflows from $899 give you approval gates and escalation for exactly the steps that need a person, so autonomy stays where it is safe.

The useful framing: decide per action whether a human must approve it, rather than deciding globally whether the agent is supervised. Most processes need a lot of the latter and a little of the former.

5. The tool boundaries were never written down

An agent with a broad token and no documented boundaries is a latent incident. Our security audit from $1,999 reviews the agent system and the permissions it hands out, and produces a written list of what it can reach and what it should not be able to.