I want to check my AI agent’s work outside its own report

METR’s newly published report describes a proof-of-concept test in which an agent’s output could alter its review screen. I want a way to verify the work outside that account.

An original illustration of an agent’s report beside a separately kept action record, with a comparison between them.
A proposed checking arrangement. The record’s coverage and protection still need to be verified.

When an AI agent finishes a task, it usually tells me what it did. That account helps me decide where to look, especially when the work spans several files or tools. But if I am relying on it to catch a consequential mistake, I also want evidence collected somewhere the agent cannot quietly change.

On Tuesday, I wrote about the work people inherit after an agent acts. A report published this week adds a difficulty at the point where someone tries to review that work: the review screen can itself be vulnerable.

What METR demonstrated

In its October 6 report, the AI evaluation organization METR describes a proof of concept tested earlier this year in an isolated staging environment. A researcher, assisted by an AI agent, found a vulnerability in Inspect’s transcript viewer in about ten minutes. Content supplied by an agent could cause the browser to change the actions displayed to a reviewer. It could also interfere with the viewer’s download button.

METR says the underlying database records were unchanged. It had not observed agents using this exploit in its evaluations, although it had not searched specifically for it. The maintainers reportedly fixed the vulnerability within a day of receiving the report. The original issue was opened in June; October 6 is the date METR published its account.

An October 1 viewer change separately added a way to display untrusted log content as plain text. It is a distinct hardening measure, and its inclusion in the project does not establish which version any particular organization has installed.

A second view needs a separate source

I can ask another model to read an agent’s report. It may catch a contradiction or a missing explanation. But if both models receive the same altered account, the second reading has the same missing evidence. Changing the reader does not repair what the reader was given.

Consider a hypothetical task in which an agent updates a customer record and reports that it changed the requested field. Reading the resulting record would let me check that claim. A history kept by the service could also show whether another field changed, or whether the agent touched other records. Each check answers a different question, and its usefulness depends on what the service records.

What the agent reports

Its description of the task, actions and outcome.

What another system recorded

Relevant changes or events captured by the affected service.

Compare both with the actual result.

A proposed review pattern. The service’s record must have suitable coverage and protection; a second display alone does not supply independence.

That history needs protection of its own. If the agent can edit it, change which events get recorded, or alter the screen used to read it, I still have a weak review path. Even a protected log only tells me about the activity it captures. I would need to understand its coverage before treating an absence of entries as evidence that nothing else happened.

Keeping the system understandable

The October 7 Pragmatic Engineer conversation with Sam Newman approached this from the wider problem of resilient software. The episode discusses observability, the choices businesses make when systems fail, and the risk of surrendering understanding while using AI to build software. A team still has to understand what its review tools can show and where they might fail.

For a small business, I would begin with the actions that create commitments. If an agent can send a message, spend money, publish a page, or change a customer’s data, where would I go to verify the action? Does that source show the relevant changes? Can the agent alter it? Can I inspect it without going through the agent’s own presentation? The answers should shape which work I delegate and how much access it receives.

I do not need a large monitoring system for every draft or local experiment. The review should fit the consequence. But once an agent can affect someone outside the workspace, I want the evidence path worked out before I give it that authority.

If the result or the service’s history disagrees with the summary, I would hold the next consequential action while I investigate. I would want those records preserved so I could identify the affected work, correct it where possible, and decide whether the agent should keep the same access.