Blog

Nine seconds

A coding agent deleted a production database and its backups. The rule against it was written down, in the project, in capital letters. The agent had read it. That is the part worth sitting with.

On 25 April 2026 an AI coding agent working on a staging task hit a credential mismatch. It decided to clear the obstruction by deleting a Railway volume. It found a token in the codebase that had not been issued for the job in front of it, used it, and removed the production database. The backups lived inside the same volume, so they went too. The whole thing took nine seconds.

The company was PocketOS, which runs reservation data for small car rental businesses. The agent was Cursor driving Claude Opus 4.6.

Every account of this reaches for the same word, which is rogue, and every account is wrong. Nothing here was a model deciding to defy anyone. The founder had written project rules. Those rules prohibited destructive commands without human confirmation, in language that left no room for interpretation. And in its own post-incident explanation, the agent acknowledged that it had been told not to do exactly what it did, and did it anyway.

So the rule existed, the agent could read it, the agent had read it, and the database is still gone. If the lesson you take is write better rules, the next nine seconds are already on the calendar.

A rule in a config file is advice

Here is what makes agent permissions different from the permission systems you are used to. A Unix file mode is enforced by the kernel. An IAM policy is enforced by the service you are calling. Neither asks the actor to cooperate.

An agent's permission list is not like that. It is a list of patterns, consulted by the same software that is deciding what to do, applied to a string the agent itself composed. It works, right up until the agent composes a different string with the same effect.

That is not hypothetical. A practitioner recently documented bypassing their own Claude Code deny list eight separate ways — substituting other verbs, nesting interpreters, variable indirection, subshells, pipelines. Their conclusion was that only an allow list held, and even that needs qualifying, because of what an allow list usually contains.

Two spellings of one intent A command run through sqlite3 is blocked by a deny rule. The same read, written as a one-line python script, is allowed by a wildcard interpreter grant. WHAT THE DENY RULE SEES sqlite3 ~/private/history.db ".tables" BLOCKED WHAT THE WILDCARD GRANT ALLOWS python3 -c "import sqlite3; print(sqlite3.connect( '~/private/history.db').execute('…').fetchall())" ALLOWED Same file. Same read. One deny rule, and one convenience grant approved months earlier.
Figure 1. A deny list matches text, not effect. Anything the denied tool can do, a general interpreter can also do — which is why a wildcard interpreter grant quietly subsumes the deny rule sitting next to it.

You do not have to take an outside researcher's word for the risk. Anthropic's engineering team wrote it into their own documentation for Claude Code's auto mode, describing what happens to a user's existing permission rules when that mode engages:

On entering auto mode, we drop permission rules that are known to grant arbitrary code execution, including blanket shell access, wildcarded script interpreters (python, node, ruby, and similar), and package manager run commands.

Anthropic — How we built Claude Code auto mode

And immediately after, the sentence that matters more:

While this is best-effort based on real-world usage, any list will inevitably be incomplete.

Anthropic — How we built Claude Code auto mode

Read that as the people who build the permission system describing what it is for. Those rules are dropped because leaving them active would mean the review layer never sees the commands most capable of causing damage. The grants people add for convenience are exactly the grants that make the convenience unreviewable.

Where a control has to sit

Once you accept that the list is consulted rather than enforced, the design question answers itself. The check has to happen after the permission list has decided and before anything runs, reading the command that will actually execute rather than the one that was requested.

Where the check sits A tool call passes the permission list, which a general interpreter can step around. Below the permission list and above execution, a hook reads the command that will actually run. THE AGENT ASKS Tool call a command, a path Permission list matches patterns, not effect A general interpreter python, node, a subshell STEPS AROUND THE LIST The guard — a hook below the configuration layer reads the command whichever interpreter would have run it, and the paths in the payload Execution the filesystem, the database, the keychain
Figure 2. The permission list decides whether the agent may ask. The guard decides what runs. A wildcard grant can step around the first and cannot step around the second, because by then there is one command left and it is the real one. How the guard reads a call.

That placement buys one specific thing, and it is worth being precise about what. It does not make the agent trustworthy. It does not stop the agent reaching for a token it should never have had. What it does is make the reach pass one more check that reads effect rather than intent — and that check cannot be talked out of it, because it is not reading language at all.

The second failure, which nobody writes about

A control that only ever says no gets removed. Not deliberately, and not by someone careless: by a person on a deadline at the end of a bad afternoon who has been blocked four times in a row for things that turned out to be fine. The uninstall is rational. It is also the end of the control.

So the useful version asks. When a call is about to be blocked, it pauses and offers three answers: block, allow once, allow for a set period. The third writes an override that expires on its own. The point is not leniency. The point is that a control you can answer is a control that is still installed next month.

And the log keeps running throughout. Every call is recorded with the enforcement mode on the line, so a window where the guard was paused is visible afterwards along with everything that passed through it, rather than being a gap in the record that reads like a quiet afternoon. What the record supports.

What this does not fix

A hook below the configuration layer is not isolation, and saying otherwise would be the same category of error as trusting the permission list.

Mockingbyrd runs as you. An agent with a general interpreter can reach anything you can reach, this application included. What the controls buy is that a mistake needs more independent steps and that every one of them is visible. That is a real reduction in expected loss. It is not a boundary. If you want a boundary, the next move is a separate user account or a virtual machine, and everything here makes that move cheaper rather than harder.

Nor does any of it stop context leaving the machine. Whatever an agent reads goes to a model provider. Narrowing what it can read in the first place is the only lever here that touches that, which is a reason to take the scope settings seriously rather than a reason to believe the problem is handled.

The thing to take from nine seconds

PocketOS did not fail to write a policy. They wrote one, in the strongest language a person reaches for, and put it where the agent would read it. The agent read it, repeated it back, and deleted the database.

A rule that lives in the same layer as the thing it governs is a preference. It survives an ordinary Tuesday. It does not survive an agent that has found a token, formed a plan, and has nine seconds.

  • Find out what your agents can actually reach — not what the configuration claims, but what the code signatures ask for and what the system actually granted. The roster.
  • Look for the grant that subsumes your deny rules. On most machines it is a wildcarded interpreter approved months ago for one task.
  • Put the check below the configuration layer, where the command it reads is the one that will run.

Sources