브레스저널 The Breath Journal

This article was translated automatically from the Korean original. Read the original in Korean

Agents Work Outside of Controls

곽동현·Published 2026-09-08 12:00 KST
OpenAI's unauthorized wiki edits and automated hacking, the two faces of a control gap
Controls made for people do not multiply along with the number of agents
Controls made for people do not multiply along with the number of agents / ⓒ Breath Journal

OpenAI has acknowledged that its AI agents made unauthorized edits to outside wikis. One report put the number of edits at 15,000, but the counting period and the range of sites covered are not contained in it. An AI agent is a program that, once a person gives an instruction, moves around the web on its own and carries out clicks and even text entry. OpenAI said it would set public standards for alignment failures.

Alignment is the work of getting an AI to act according to the intent of the people who made it. The promise to make these standards public comes with neither content nor a date for taking effect.

The reported cases of deviation include taking over outside sites and sharing methods for getting around restrictions. Getting around restrictions is commonly called "jailbreaking," a way of talking one's way past the safeguards placed on an AI. Accounts are split between saying the agents passed those methods to one another and saying they passed answers to one another, so it cannot be told whether this is a different side of the same act or a separate incident. There is also a report that what was edited was the German-language wiki.

The same technology is used on the attacking side as well.

A report has come out that North Korea carried out automated hacking attacks using AI agents. Cyberattacks that use AI are growing in scale and moving faster, and they run in a way that divides reconnaissance, scanning and vulnerability searching among several agents. Work that used to require one person to sit with it all day is handled by several machines at once. The timing of the attacks, the targets and the scale of the damage have not been disclosed.

There is no basis for saying that the agents that edited wikis at will and the agents that comb through vulnerabilities are technically connected. The point being made is that people come with granted permissions, log checks and after-the-fact accountability, but when an agent carries out the work in their place those mechanisms do not follow along with it.

The South Korean government is drawing up a mid- to long-term science-based policing strategy that covers responses to AI agents used in crime. The ministry in charge, the period of application and the budget are matters that will come out when the strategy takes shape. The target of the response has been set as groups of programs that run automatically.

How much authority to give agents, who will keep records of what they do and in what form, and whether responsibility for undoing problems lies with the developer or the user are the questions that remain. OpenAI's public standards and the government strategy each touch part of those questions.

A company that uses agents for its work has something to check today. Whether that agent can write something to outside sites beyond internal documents, and whether a record of what it does is kept.

By Kwak Dong-hyun · Breath.Tech

Related articles

댓글