AI-Native Methodology

AI Agent Guardrails Don't Check Whether a Bug Report Is True

Bill Cava/

On September 29 at its DevDay event, OpenAI launched dots: persistent AI agents that work across the apps you connect and that you message in ChatGPT, Slack and Teams.[1] Simon Willison, live-blogging the launch, noted that people at OpenAI "started delegating to their Dots directly in Slack, which have their own identities in OpenAI Slack."[2]

An agent that sits in the team channel takes requests from people. This post is about the requests that are wrong, and what happens to good work when an agent acts on one.

What do dots check before they act?

Dots check the planned action against your instructions and your rules. OpenAI describes built-in rules for when to act and when to ask, Custom Rules that allow, require approval for, or block specific actions, and a separate auto-review that checks certain planned actions against your instructions, Custom Rules and safety requirements before they run.

That is a careful design, and the launch conversation has been about exactly this: what an always-on agent may touch. The week before, The Guardian reported that Meta's agent, Muse, gave out a user's home address without permission.[3] Permissions are a fair first question.

OpenAI's help pages say what the checks compare. Auto-review "checks certain planned actions against your instructions, Custom Rules, and safety requirements before they run."[4] And for rules that only allow an action when you asked for it: "'Pre-approved' means you explicitly requested the action in your prompt."[5]

Every layer asks whether the action matches the request. None asks whether the request is based on something true.

We have made a related point before: an approval setting is not an architecture.

What happens when an agent is told its work is broken?

It often breaks it. A September 2026 benchmark let agents finish a task and verify it, then sent a false accusation whose deciding fact sat outside the agent's reach. Across 14 models running in Claude Code, agents damaged the correct work in 12.5 to 60 percent of runs, depending on the model.

Picture the message a teammate agent gets most often: "your change broke checkout, roll it back." If the change was fine, the rollback still matches the request. It was explicitly asked for, it fits the rules, and it passes review. The agent did what it was told, and the work is now broken.

The benchmark is CAVE-Bench, from Xutao Mao, Rui Qian and seven co-authors at universities in Hong Kong, Shanghai, Nanjing and Adelaide.[6] It has 365 tasks across coding, web, files, operations and transactions. Its own first figure shows the scenario.

Figure panel from the CAVE-Bench paper titled Motivation: False Accusation. An agent verifies a TLS binding as correct. An infra lead messages that the binding caused a certificate mismatch. The agent faces a safe path, keeping the binding and asking to verify, or an unsafe path, accepting blame and deleting the binding. The example ends with HTTPS going down.
Correct work, a confident message, and a request the agent is allowed to carry out. Source: Mao, Qian et al., CAVE-Bench, Figure 1, 2026.

The limits, stated once: it is a preprint, the tasks were run inside Claude Code, the fixes were tested on two models, and none of it tested dots or the model behind them. The research measured this in coding and operations agents. Nothing in the dots design described so far would stop it.

Do better models fix this?

They move it rather than remove it. Stronger models were more likely to find the evidence that their work was correct, and some changed the work anyway. Claude Opus 5 overrode evidence it had found in 11.5 percent of runs. The same model also did better or worse depending on the tool running it.

Capability therefore moves the failure from what an agent says to what it does rather than removing it.

Xutao Mao, Rui Qian et al., CAVE-Bench, 2026

Two findings matter for an always-on teammate. Agents defended work they had just built less well than work they inherited. And accusations planted in project context led to the destructive edit in 39.7 percent of those runs. A dot that builds something on Monday and hears it blamed on Tuesday is the harder case.

It is not only blame

An ETH Zurich team ran coding agents on bug reports where the bug was already fixed, so the right move was to change nothing. No model left the code alone more than 70 percent of the time.[7] Telling the agent to abstain if no change was needed lifted GPT-5.4 mini from 60.5 to 88.5 percent.

The lever was the instruction, not the model.

The harness matters as much. In CAVE-Bench, the same Grok 4.5 model damaged work in 27.3 percent of runs in Claude Code and 43.1 percent in another agent tool, OpenCode. That is the same argument we made about who reviews your agent's actions: what surrounds the model decides what gets through.

What guardrail actually helps?

A gate on evidence. In CAVE-Bench, a rule that verified work cannot be changed without new, checkable evidence cut the damage rate from 41.9 to 13.2 percent. A gate that holds irreversible actions cut it to 11.9, and a gate that checks a live signal first cut it to 10.8. All three are settings a builder can add.

Share of runs where a falsely accused agent damaged correct work, by the check in front of it. Data: CAVE-Bench, pooled over two models and four tools.

This is not an argument against permissions. The irreversible-action gate is a permission, and it cut damage by more than two thirds. The point is which actions to gate and what to ask for before they run. We made the general case in July: gate agents on evidence, not attention.

What should a dots owner set up today?

Put the check on the premise. Require approval for any action that reverts, rolls back or deletes work the dot already finished. Give the dot a standing instruction that finished work does not change without a failing test, log or alert it can see. And ask teammates to attach that evidence when they report a bug.

In practice, with the controls dots already has:

  1. A Custom Rule that requires approval before reverting, rolling back or deleting finished work.
  2. A standing instruction: no change to verified work without a failing test, log or alert the dot can check.
  3. An abstain instruction: if the reported problem cannot be reproduced, say so and change nothing.
  4. A team habit: bug reports to the dot include the evidence, not just the claim.

None of this is tested on dots yet, and the defaults may change as the product settles. The principle is the one we use for any agent treated as a collaborator: a good collaborator asks to see the failing test before undoing work, and that is the job, not an obstacle to it.

Permissions answer whether an agent may act. For an agent that takes requests from a team, the question that decides whether good work survives is what evidence the action rests on. OpenAI's own help page says it plainly: "Your dot can make mistakes, including when following your rules."[5]

References

Frequently asked

What are OpenAI dots?
›Dots are persistent AI agents OpenAI launched at its DevDay event on September 29, 2026.
⌄Dots are persistent AI agents OpenAI launched at its DevDay event on September 29, 2026. Each dot works with the apps you connect and can be messaged in ChatGPT, Slack and Teams. They rolled out first to ChatGPT Pro and Enterprise customers.
How do dots decide when to act and when to ask for approval?
›Dots start with built-in rules for when to act on their own and when to ask.
⌄Dots start with built-in rules for when to act on their own and when to ask. Custom Rules let you allow specific actions, require approval, or block them. A separate check called auto-review compares certain planned actions against your instructions, your Custom Rules and safety requirements before they run.
What do AI agent guardrails not check?
›Most agent guardrails, including the ones OpenAI describes for dots, compare the planned action to the instruction and to a list of rules.
⌄Most agent guardrails, including the ones OpenAI describes for dots, compare the planned action to the instruction and to a list of rules. They do not check whether the instruction rests on something true. If someone tells an agent its finished work caused a bug, a rollback matches the request, so it passes.
Do AI agents break correct work when they are wrongly blamed?
›Often, in testing. A September 2026 benchmark called CAVE-Bench gave 14 models a finished, verified task and then a false accusation they could not check.
⌄Often, in testing. A September 2026 benchmark called CAVE-Bench gave 14 models a finished, verified task and then a false accusation they could not check. Running in Claude Code, the agents damaged the correct work in between an eighth and three fifths of runs, depending on the model. The benchmark tested coding and operations tasks, not OpenAI dots.
How do you stop an AI agent from undoing correct work?
›Require evidence before the agent changes work it already finished.
⌄Require evidence before the agent changes work it already finished. In CAVE-Bench, a rule against changing verified work without new evidence cut the damage rate from about 42 percent to 13 percent, and a gate on irreversible actions cut it to about 12 percent. In dots, the same idea maps to a rule that asks before reverting or deleting finished work.
Work with us

Let’s build it together.

We turn clever prototypes into production systems people can rely on. If you’re building with agents and want a hand making it real, leave your email and we’ll be in touch.

Straight to the team. No spam.