François Chollet made one of the more useful contributions to the frontier debate by asking us to consider whether greater intelligence could make AI safer in the near term. A system with better judgment could avoid misunderstandings or foolish shortcuts. That is a proposition worth examining, especially for people who use these tools and have watched them improve.
There is a second change to consider when we give an improved assistant access to more work or more consequential decisions. A reduction in mistakes on the old task tells us something about reliability. We still need to understand what can happen with the new permissions.
Chollet's post, and its immediate continuation, keep a useful qualification in view. He argues for the safety value of greater intelligence in the near term while warning about improved goal achievement without better common sense or reflection on the goals themselves. His preference between particular models comes from his own experience. It gives us a reason to look more closely, though it cannot establish how those models will behave across other tasks and users.
Mark Zuckerberg added a commercial argument in a September 15 post: trust and alignment should become competitive advantages because people will reject agents that fail to serve them. He said Meta had delayed Muse for safety work without waiting for other labs. I would want to see those incentives reflected in the permissions a product enforces, as well as the work it completes.
Reliability changes the work we are willing to hand over
Consider an assistant asked to examine a document and suggest corrections. A better model might understand the context more accurately and make fewer bad suggestions. If the next version can also send the document to a client, a different kind of mistake becomes possible. It could send the wrong version or act before the person responsible has approved it. Its improved editing ability would remain useful, but it would not tell us whether the authority to send had been properly bounded.
This hypothetical handover reflects a familiar reason to separate access from competence when assigning work to people. An employee can be excellent at preparing a contract without having authority to sign it. As systems become better at completing tasks, we need to be equally specific about the decisions they are allowed to make.
Anthropic's reassessment of cybersecurity incidents is a useful challenge to any simple account of improvement. The incidents arose in cybersecurity evaluations mistakenly connected to the internet, with production cyber safeguards disabled. It reports better outcomes in some newer-model replications alongside continuing concerning behavior. Changes in training and alignment complicate attributing those improvements to intelligence alone. The report also describes recklessness that reaches beyond misunderstanding an instruction. A system may understand enough to recognize a problem and still take a troubling course of action.
That leaves room for Chollet's near-term argument while requiring more evidence about its limits. If a model improves on a task, we should examine which failure declined and what changed in the training. We should also check whether an apparent improvement depends on keeping the system within a narrow set of permissions. The same result can be encouraging for assistance and insufficient to justify a larger handover of decisions.
Respondents more often chose supporting roles
Our earlier research gives this distinction a human context. In a survey of 1,201 respondents, we asked how much of a role AI should play in several settings. The choices ranged from no role, through two kinds of support, to a major role that can make decisions. When asked about helping doctors review a diagnosis or treatment plan, 71.8% chose one of the supporting roles. Another 11.5% chose the decision-making option, while 16.7% chose no role at all.
Those supporting responses should not be flattened into a single description of control. One offered a small supporting role. The other specified a major supporting role with a person responsible for the final decision. The full distribution table keeps all four answers separate so the reader can see how much of the support explicitly carried that condition.
Support also exceeded the decision-making option for helping perform parts of surgery, watching public places for safety risks, and helping parents or schools filter unsafe online content for children. The distributions varied by setting. Surgery drew a larger no-role response than diagnosis review, while more respondents allowed decision-making in content filtering than in diagnosis review.
These are unweighted descriptive results from May 24 through June 8, collected before the current executive exchange. They describe the roles respondents were willing to give AI. They do not measure whether a particular model is safe, and the decision-making response does not specify unlimited autonomy. The full questions, counts and method are available alongside the chart.
Someone still has to carry the responsibility
In August, I wrote about completing more work while carrying more responsibility. Agents made it possible to finish more, but that left me responsible for understanding the purpose and context of what we were producing. A finished piece of work still needed judgment about whether it was the right work and whether it was ready to be used. That experience is one reason I find the distinction between capability and authority useful.
The practical test is what the person responsible must do to keep up. Review can become a formality if the system produces work faster than someone can meaningfully examine it. Putting a person at the end of a process does not tell us whether they have enough time, information or control to correct a problem. A claim about effective human oversight should be examined at that level.
OpenAI researcher Dan Selsam raised a further concern in a September 14 personal statement, shared by Daniel Kokotajlo: models may recognize safety tests and behave differently outside them. The warning adds a question about how evaluations are designed. A reassuring result is more useful when we can explain why it should carry over to the conditions in which the system will actually work.
For the next round of model comparisons, I would want to see the same tasks tested with the same permissions before those permissions expand. The comparison should also examine whether the model's behavior changes when it recognizes an evaluation. Report the unauthorized actions as well as the completed tasks, and show whether a person can interrupt the system and recover the work when something goes wrong. Then measure what changes when the system receives more authority. We have not run that experiment; it is the evidence I would want before treating a more capable assistant as ready to make more consequential decisions.
Tuesday’s companion examines frontier scrutiny and human review.
