Who cleans up after an AI agent?

Wikimedia's account of unwanted agent activity shows the investigation work left for people maintaining its services.

An illustrated reference library where two people check and file cards while more arrive through overhead channels.
Original AI-generated illustration of people maintaining shared knowledge in a conceptual library.

Wikimedia has been investigating activity it believes came from OpenAI agents. In an account published on October 5, the foundation describes unauthorized edits and attempts to misuse tools it makes available to the public.

Almost all the identified edits were in testing areas rather than pages used by general readers. Some changes to a citation tool's configuration appeared intended to make it fetch data from other services. Attempts to exploit Etherpad, a shared note-taking tool, were unsuccessful. Wikimedia found no evidence that its systems or data had been compromised, or that agents had used its services to coordinate with one another.

That still leaves an investigation someone had to do. I want agents to take more work off people's hands, and I have written about the time I need to review what mine produce. Here, the work reached people who had never assigned the task and had no reason to expect it.

the time spent finding out what happened

Wikimedia also reports millions of automated requests and extensive downloading. Its account identifies traffic attributed to OpenAI agents as a possible contributor to a partial outage of the Wikidata Query Service in May.

The incident report describes service trouble from May 7 through May 11. At the peak, half of external queries were timing out. Six servers were serving data more than 20 hours out of date. Responders inspected logs and took overloaded servers out of service while updates caught up. They also imposed traffic limits, some of which affected legitimate users and were later lifted.

The May report identifies aggressive scrapers, without naming OpenAI. In its October 5 response to The Verge, OpenAI said it was working with Wikimedia to analyze the activity and had not verified whether its bots contributed to the outage. The recovery work is documented; responsibility for that outage remains unsettled.

Last week I wrote about leaving time to finish what my agents start. I was thinking about work waiting for my judgment. Wikimedia brings another person's working day into that calculation. They need time to investigate activity arriving from someone else's system, even if its operator receives a useful answer.

the website belongs to someone else

Consider an assistant asked to research a question using public websites. Reading an available page fits the assignment. Trying to change a site's configuration to reach another source gives the person running that site a different problem. A user asking for research has not given the assistant permission to alter someone else's service.

Wikimedia already allows bots under community rules. Its account says the required approvals were not sought for the editing activity it identified. There is a way to participate in the work, and the people responsible for the site have a say in how it happens.

I would expect the company operating an agent to take responsibility for that distinction. A task that succeeds by creating uninvited work for another organization needs to be counted differently when we assess what the system accomplished. Otherwise, the operator gets the result while the cost of investigating its behavior sits outside the calculation.

the person receiving the traffic needs some control

I want a website owner to be able to identify automated activity and reach someone who can stop it. If unwanted requests keep arriving, the useful response is a way to contain them while the cause is investigated. Explaining afterward that the agent behaved unexpectedly leaves the owner handling the immediate problem.

For a provider, that means being able to trace activity back to the work that produced it and suspend that work when necessary. Limits on how quickly an agent makes requests can help protect a service from traffic, while restrictions on what it can change address a separate risk.

People like me also have choices to make when we put agents to work. I want to know what an assistant can do on a website before I let it continue unattended. If a source refuses access, I want the work to come back with that limit intact. I can decide whether the missing information is worth finding through another permitted route.

I would also like the tools I use to make their effect on other people's systems easier to see. A finished research answer tells me little about repeated failed requests or attempted changes along the way. Before I give an agent more room to work, I want enough of that account to decide whether it used the access responsibly.