For a long time in human history, the search has been on for the "right person". A country hopes to encounter a wise and just ruler, an organization hopes to have a loyal and dedicated manager, a business system hopes to find a trustworthy agent, and a family hopes to entrust important matters to a reliable person. As long as this person is intelligent enough, kind enough, and self-disciplined enough, many complex problems seem to be naturally resolved.
This way of thinking hasn't disappeared—it's just that today, we've swapped in AI for that anticipated "perfect subject." We want models to be smarter and better aligned. We want them to accurately understand instructions, avoid hallucinations, resist manipulation, and know what should and shouldn't be done. As AI agents begin to gain operational capabilities over tools, accounts, APIs, funds, and even physical devices, these expectations grow ever stronger, gradually crystallizing into what looks like a perfectly reasonable goal: we need an agent that is safe enough, reliable enough, and obedient enough.
But if we take a longer time scale, this goal actually hides a very old governance logic - we are once again trying to solve the problem of "bad outcomes" by searching for "good subjects". And modern institutions are what they are today precisely because humans have gradually abandoned this illusion.
1. The most natural answer to early governance is to find a “good person.”
When a society is small in scale, building order on personal virtue is not without reason. A tribe may consist of only a few dozen people, a commercial organization may have just a handful of core members, and the decision-making chain may span only two or three layers, with highly repetitive, long-term relationships among individuals. In such an environment, reputation, morality, loyalty, and personal judgment are themselves effective governance mechanisms: if leaders are sufficiently restrained, many rules need not be written down; if agents are sufficiently loyal, many permissions need not be strictly delineated; and when participants know one another well, many risks can be absorbed through social relationships themselves.
Throughout a long period in human history, when organizations faced governance dilemmas, the most intuitive response was not to redesign the system, but to search for better people - hoping for wiser monarchs, more honest officials, more trustworthy merchants, and more loyal executors. This is not foolish, the real problem lies in that it cannot be scaled. Because once a society expands, it will inevitably collide with an unavoidable fact:
A system cannot assume that all participants will always have accurate information, proper motivations, and sound judgment.
People can become fatigued, misinterpret, be tempted by benefits, and make incorrect choices under pressure. Even if a person's character never changes, the information they possess may be incomplete, and their understanding of reality may be incorrect. Therefore, "finding a good person" is no longer sufficient to address governance issues, and humans begin to ask another question: if we cannot guarantee that every person in power is a good person, can we at least ensure that the mistakes of bad people, wrong people, and ordinary people do not easily destroy the entire system? This is where institutions truly begin to mature.
Second, the breakthrough of modern systems lies not in finding better people, but in acknowledging that people are imperfect.
The most crucial, yet often overlooked, ideological shift in modern governance is transforming "imperfection" from an occasional anomaly into a design premise. Power requires checks and balances, not because we are certain that a specific individual will abuse their power, but because institutions cannot be based on the assumption that "they will never abuse it." Companies need financial approvals, segregation of duties, and internal controls, not because employees are assumed to be untrustworthy, but because mature organizations cannot rely on any single entity to always be correct. Aviation systems establish cross-checks, proceduralized operations, and redundant designs, not because pilots will definitely make mistakes, but because pilots may make mistakes.
A profound shift has occurred here: the object of governance is no longer just the "bad guys", but a more universal subject - those who make mistakes. There is an essential difference between the two. If all risks come from malice, we only need to identify the bad guys; but in the real world, the most difficult risks to deal with often come from people without malicious intent: cognitive biases, errors in judgment, operational mistakes, outdated information, lack of context, collaboration errors, and decisions that are completely reasonable locally but lead to severe consequences overall.
It is for this reason that mature systems have never been concerned solely with "who is trustworthy," but must also answer the question "what happens if this trustworthy person makes a wrong judgment today." This is the true dividing line between good people governing and good systems: the former attempts to increase the probability of correct decisions, while the latter further inquires -
When the subject is incorrect, can the system still remain within acceptable limits?
Thirdly, AI Agent is experiencing a new round of "good person politics"
It's interesting that with the emergence of AI agents, we seem to have returned to a familiar starting point: we're looking for a "good" agent. Many of today's discussions about agent safety still revolve around the subject, focusing on making models smarter, more aligned, more user-friendly, less prone to hallucinations, better at recognizing malicious prompts, and clearer about what not to do. These directions are certainly important and should continue.
The issue is that when an Agent evolves from a system that only generates content to an entity that can actually change the external world, simply improving the entity itself is no longer sufficient. This is because a easily overlooked change has occurred: errors have acquired the ability to be executed on a large scale for the first time. If a person misunderstands a sentence, they may do one thing wrong; if an Agent misunderstands a sentence, it may invoke dozens of tools, modify hundreds of objects, and spread the same error to tens of thousands of targets in just a few seconds.
Human error is subject to natural physical limitations - the need to sleep, move, wait, and limited attention and time, as well as a limited number of objects that can be affected simultaneously. Agents do not have these constraints. So, the real concern is not just "making mistakes", as humans have always made mistakes; the real difference is:
We are starting to mass-produce for the first time an imperfect entity that can replicate errors at high speed, automatically amplify errors, and continue to execute errors.
If we continue to follow the logic of "as long as the main body is trained better, it will be fine," we are actually giving more and more real power to a main body that cannot be proven to be always correct. This is not as far removed from ancient societies' expectation of a forever wise monarch as one might imagine.
The issue has never been "whether the Agent will make mistakes," but rather "whether mistakes have the power to become reality"
It is necessary to distinguish between two issues that are often confused. The first is: Will Agents make mistakes? The answer is almost certainly yes. Any subject that relies on limited information, probabilistic judgments, complex contexts, and external environments cannot guarantee absolute correctness; moreover, "correctness" itself often depends on specific business contexts, rather than being an absolute standard independent of the environment.
The second question is the true variable in governance: does a mistaken judgment have the capacity to directly alter reality? A model incorrectly concluding that it should execute a transfer is not the same security incident as it actually completing one. An agent wrongly determining that a database should be deleted is not the same problem as the database being deleted. An automated system producing an erroneous plan is not on the same level as that plan crossing every boundary and reaching the production environment.
Therefore, we may need to re-examine the boundaries of AI Safety. Safety is not just about reducing incorrect judgments, but also about a mature system's ability to manage the real-world consequences that arise from incorrect judgments.
Risk not only lies in whether the subject makes a mistake, but also in how much power the mistake has to be executed.
This is also why "a better Agent" and "a safer system" are not the same concept. You can have an Agent with a very high average accuracy rate, yet still construct a dangerous system - as long as one mistake is enough to produce irreversible, large-scale consequences; conversely, an imperfect Agent can also be deployed in a highly reliable system, provided its errors are structurally limited to a tolerable range. The aviation, finance, power, and industrial control industries have long accepted similar facts: no one requires a component to never fail, and the real problem engineering solves is whether the entire system remains safe when a component fails. AI Agents will ultimately be no exception.
5. True modernization is the transition from "believing in the subject" to "limiting power"
A mature system does not mean that trust is no longer needed between people, on the contrary, a society without trust can hardly function. What has really changed is that trust no longer equals unconditional authorization. Modern companies can highly trust their CFOs, but they will not cancel financial controls because of this; banks can trust their core employees, yet they still establish dual reviews, quota controls, anomaly monitoring, and separation of duties; a country can produce leaders through elections, but it will not assume that all power does not require constraints after the election.
The underlying logic is crucial: whether a subject is trustworthy and whether it should have unlimited power are two completely different issues. Yet, today's discussions about Agents often blur these lines - because it has passed authentication, it is allowed to call tools; because it is an officially deployed Agent, it is granted long-term credentials; because it uses a security-trained model, its behavior is assumed to be trustworthy by default; because the Prompt comes from a legitimate user, its subsequent actions are assumed to naturally inherit the user's intent.
Six, good systems do not eliminate mistakes, but rather make mistakes tolerable
It's easy to swing to the other extreme: if the subject is unreliable, should we impose more and more restrictions on the Agent until it can't do anything? Of course not. The only way to prevent all errors is to prevent a multitude of correct actions at the same time; absolute safety often means absolute loss of capability.
The truly difficult part of system design is not eliminating risks, but establishing an acceptable risk structure: what is allowed, what is restricted, which mistakes can be recovered from, which mistakes must be prevented before they occur, under what circumstances can things be automated, under what circumstances do new participants need to be added, and under what circumstances the system should retain the ability to refuse to continue even if all identities and permissions are legitimate. The goal of a good system is not to create a world where mistakes never happen, but to give the system the ability to allow mistakes to occur without allowing them to expand indefinitely.
This is also a key reason why modern society is able to form complex collaborations. We are not brave enough to sign contracts because we believe that all business entities will never default; on the contrary, it is because contracts, liabilities, laws, insurance, audits, guarantees, and the judiciary form a complete set of systems to deal with "possible defaults" that we dare to expand our cooperation to include strangers.
Institutions are not the opposite of trust; rather, good institutions are the foundational infrastructure that enables large-scale trust to exist.
This also holds true in the Agent world: only when companies no longer need to trust that Agent is always correct can they truly feel at ease handing over increasingly important tasks to Agent.
7. What AI truly needs may not be more ethics, but rather more institutions
In recent years, we have extensively discussed AI ethics, including issues such as fairness, transparency, bias, privacy, accountability, and interpretability, all of which are extremely important. However, as Agents begin to take action, a new issue is becoming increasingly prominent - how can ethical principles be translated into real-world actions?
But when agents begin to possess real execution capabilities, we cannot forever stay in the stage of value statements. Because reality will ultimately not ask about an agent's values, it will only leave results: were the funds transferred, were the servers deleted, was the code deployed, were the devices started, was data leaked, and was the asset status changed.
This is also why the Agent era will eventually have its own system engineering: translating abstract principles into boundaries for action, translating value judgments into conditions for execution, and translating "should" into "under what circumstances can something truly happen". This is not a negation of ethics, but rather the way ethics enter the real world.
The most important question for the future may not be "whether AI is like humans," but rather "whether we will govern it in the same way we govern humans."
For decades, a question has lingered around artificial intelligence: will machines think like humans. But in the era of Agents, the truly pressing issue may not be this one. Even if it never thinks like a human, it has already begun to possess some capabilities that only human actors once had: accepting goals, interpreting tasks, selecting tools, devising steps, invoking resources, altering external states, collaborating with other agents, and leaving consequences in the real world.
Once a system possesses these capabilities, governance issues arise. We don't need to first prove whether an Agent has free will or resolve whether a machine has true consciousness to discuss what degree of agency it should have. Companies are not biological humans, yet modern society has designed independent rights, responsibilities, and constraints for them; governments are not natural persons, but political institutions still expend great effort to restrict how they exercise power. Agents may ultimately become a new type of actor: they may not possess full personhood, but they will have sufficient agency to produce real-world consequences.
And once that happens, we must confront an age-old question:
Any entity capable of altering reality should be empowered in a manner that balances its potential benefits with necessary restrictions to prevent abuse and ensure accountability.
This is no longer just an AI issue, it's an institutional issue.
9. AI Agents must also undergo modernization
From this perspective, the position AI Agent is in today is similar to an early stage in the history of human systems. We are still highly focused on the subject itself: hoping it is intelligent, kind, and loyal, accurately understanding our intentions, and trying to make it a "good subject" worthy of being entrusted with power through training, rules, and value alignment. These efforts are all worth continuing.
But the real turning point may come on another day - when we begin to accept a more mundane, yet more mature fact: Agent will never be perfect. It will make mistakes, misinterpret, face unprecedented situations, receive incomplete information, be affected by adversarial inputs, and produce overall incorrect results based on locally correct logic; and as the system becomes increasingly complex, we may never be able to exhaustively anticipate all its possible behaviors.
This does not mean that Agent cannot enter production, on the contrary, it means that we can finally stop waiting for the "perfect Agent" that will never appear, and instead solve a problem that can truly be engineered: how to enable an imperfect entity to participate safely in reality. This may be the true modernization of AI Agent - not making machines finally become infallible good people, but making our systems no longer require any entity to be infallible good people.
In conclusion, civilization is not complex because people have become perfect.
Looking back at human society, one interesting fact stands out. Modern business is able to exist not because entrepreneurs are no longer greedy; modern governments are able to function not because those in power no longer make mistakes; and modern finance is able to handle massive assets not because every person in a bank is always reliable. Modern civilization has been able to become so complex largely because we have gradually learned one thing: not to require every entity within a system to be perfect.
We create contracts, laws, audits, insurance, separation of duties, checks and balances, and accountability systems, not because we have lost faith in people, but because we have finally come to realize that people are trustworthy, yet people make mistakes; people can hold power, but power needs boundaries; people can act autonomously, but autonomy should not mean unlimited consequences. This is the process of transitioning from "trustworthy individuals" to "trustworthy structures".
Today, as AI agents begin to penetrate funding, software, cloud systems, corporate operations, and even physical devices, we are once again faced with the same choice. We can continue to seek a smarter, more loyal, more aligned, and infallible "good agent"; or we can acknowledge that the truly mature era of agents will not be built on perfect agents, but on a more ancient and reliable civilizational wisdom:
Do not rely on any single entity to always be correct for the safety of the world.
The true progress of civilization may never be about finding enough good people, but rather about establishing a system where the world can still function even with ordinary, flawed, and even bad people.
AI agents must also undergo this modernization process.
