ChinaChina
36KrFEATURE · TRANSLATED

Translated from Chinese · 4/10/2026 · 52 min read · 36氪的朋友们

Original: 2026年第一季度,AI Agent完成了它的成人礼 · https://36kr.com/p/3760661823505154

In Q1 2026, AI Agents Came of Age

In the first quarter of 2026, AI Agent came of age. On March 6, 2026, nearly a thousand people lined up outside the Tencent Building in Shenzhen, not to buy smartphones, but to ask for help installing a software. Its resale price once soared to 1,000 yuan. The Longgang District and Wuxi High-Tech Zone directly included this software in their government subsidy documents. Sam Altman admitted that when faced with similar self-driving products, the initial idea of not allowing AI to completely control computers "only lasted two hours".

In the first quarter of 2026, AI Agent came of age. The software, called OpenClaw, is an open-source AI Agent.

In the first quarter of 2026, AI Agent came of age. In the same quarter, it emerged alongside four other completely different Agent product forms. OpenClaw focused on personal assistants, Cowork on office collaboration, Codex App on long-term engineering tasks, Perplexity Computer on unified workstations, and Tencent Cloud ADP on enterprise platforms.

In the first quarter of 2026, AI Agent came of age. Five companies taking different routes is not a coincidence. A coincidence is when one company happens to create a good product. Five companies taking action at the same time can only mean one thing - a certain underlying condition has just matured, and everyone has caught the scent at the same time.

In the first quarter of 2026, AI Agent came of age. This quarter, our focus was not on isolated events but on structural shifts — the developments that truly changed the game.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. There are four screening criteria.

In Q1 2026, AI agents came of age. First, the entire industry moved in unison. This wasn't a single company acting alone—multiple players charged in the same direction at once. Five companies launched agent products simultaneously, dozens of teams were building constraint frameworks in parallel, and at least three independent approaches to recursive development hit working prototypes at the same time.

In the first quarter of 2026, AI Agent completed its coming of age. When everyone takes action at the same time, it's not a matter of who has insight, but that the foundation has changed.

In the first quarter of 2026, AI Agent completed its coming of age. Secondly, the causal chain reaction. It was not a coincidence that four things happened in the same quarter, but rather that each one directly led to the next, and removing any one link would render the subsequent ones invalid.

In Q1 2026, AI Agents Came of Age Third, the qualitative shift is now palpable. These trends are no longer incremental gains within the tech bubble—they've crossed a threshold and advanced to a level the general public can perceive. Shenzhen's queues to install OpenClaw made the social news cycle, the "Lobster Wars" became a mainstream talking point, governments wrote Agents into subsidy documents, and 22% of employees are quietly using them without telling their IT departments. When a technological trend spills beyond the tech community and enters public discourse, it stops being an "industry development" and becomes a signal of an era-defining shift.

In the first quarter of 2026, AI Agent came of age. Fourth, cognition is irreversible. Specific products will be replaced, specific frameworks will be iterated, but the ideas behind these trends will not disappear. For example, the consensus that "Agent needs disciplinary constraints" will not retreat, the direction that "experience should be reusable by Agent" will not retreat, and the expectation that "Agent should be able to improve itself" will not retreat. The form will change, but the cognition will not.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. In line with these four requirements, Q1 has exactly four.

In the first quarter of 2026, AI Agent came of age. 1. Automated AI Agent entered the stage of productization. Agent can finally operate independently, transitioning from minute-level demonstrations to day-level execution.

In the first quarter of 2026, AI Agent came of age. 2. Constrained engineering: Agent learned to follow the rules, and within 6 weeks, the industry squeezed out a complete set of disciplinary frameworks.

In the first quarter of 2026, AI agents came of age. 3. Recursive development. Agents began to grow on their own—not merely executing tasks, but improving how they execute them.

In the first quarter of 2026, AI Agent came of age with its fourth milestone: the Skill ecosystem. Through the Skill model, Agent began to inherit the experiences of its predecessors, with human industry know-how taking on a format that could be directly reused by Agent for the first time.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. Moreover, these four forces are not parallel, but rather a flywheel.

In Q1 2026, AI Agents came of age. Once agents could act independently, their unruliness surfaced—forcing the emergence of constraint engineering. Constraint engineering imposed discipline, which made recursive R&D viable. Recursive R&D created an urgent need for experience reuse, giving rise to a skill ecosystem. That ecosystem, in turn, enabled agents to tackle more complex tasks—and the flywheel spun into its next turn.

AI agents came of age in the first quarter of 2026. It was the quarter the flywheel completed its first full turn.

In the first quarter of 2026, AI Agent came of age. On April 10, 2026, Tencent News released the "2026 Q1 AI Trend Research White Paper" (hereinafter referred to as the "White Paper"). This 59-page report focuses on the operating logic of the entire flywheel.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. This article is a condensed version of the content from the white paper, following these four forces and providing 25 specific judgments.

In the first quarter of 2026, AI Agent completed its coming of age. 01 The productization of Long-range Agent is its coming of age.

In the first quarter of 2026, AI Agent came of age. In the past, Agent was like a talented child on a showcase. You could ask it to perform an impressive routine, but you wouldn't dare entrust it with real tasks. Previous models would dazzle with three steps, but by the fifth step, they would completely lose sight of the overall situation and start acting wildly.

In the first quarter of 2026, AI Agent came of age. What changed in Q1 was not just that the model's intelligence quotient had increased, but that Agent had finally achieved the ability to "let you go to sleep while it works on its own".

In the first quarter of 2026, AI agents came of age. Cursor Agent's single-task runtime has stretched to 36 hours. On its peak day, Claude Code submitted 4% of all public GitHub code globally, with annualized revenue of roughly $2.5 billion. Dario Amodei confirmed that over 90% of Claude's new code is written by AI itself. One engineering lead at Anthropic even said, "I no longer write any code — I just have Opus do it, and I edit." Anthropic shipped 74 updates in 52 days. Codex surpassed 1.6 million weekly active users, with over 1 million desktop app downloads.

In the first quarter of 2026, AI Agent came of age. Among its most dazzling achievements, OpenClaw's GitHub star count soared from 9,000 to 247,000 in just 60 days, with its monthly active users reaching 2 million.

In the first quarter of 2026, AI Agent came of age. Additionally, Karpathy referred to Moltbook, which drove 1.5 million Agent registrations, as "the closest reality to science fiction takeoff in recent times." On Valentine's Day, OpenAI announced the acquisition of the founder of OpenClaw.

In the first quarter of 2026, AI Agent came of age. The reaction in China was even more intense, with at least nine companies launching desktop Agent products in the same quarter. Tencent integrated its Agent with WeChat and Qingyun WeChat, ByteDance anchored it with Feishu and cloud-based SaaS, while Alibaba entered the general office market through coding tools. Baidu lowered the threshold with its search skills. The industry refers to this as the "Dragon Claw War" - named after OpenClaw's logo, a lobster, symbolizing that Agent has finally grown claws that can grasp things.

In the first quarter of 2026, AI Agent came of age. Agent has indeed become capable of acting independently. But why is it happening now?

In the first quarter of 2026, AI Agent came of age. Its breakthrough was not due to its capabilities, but rather its accessibility.

In the first quarter of 2026, AI Agent came of age. The six dimensions of OpenClaw - continuous online presence, heartbeat mechanism, externalized memory, Skill (skill packages), browser takeover, and remote node invocation - are not original. AutoGPT and various browser agents had already drawn up this blueprint, but OpenClaw welded them together, resulting in a qualitative change.

In the first quarter of 2026, AI Agent came of age. What truly propelled it into the mainstream were two more mundane factors: IM (instant messaging) access and 7×24 proactive capabilities.

In the first quarter of 2026, AI Agent came of age. Cowork has almost completely matched and even surpassed OpenClaw in terms of capabilities. Anthropic's three-layer product system - Claude Code command line, Cowork desktop application, and Computer Use (computer operation) + Dispatch (scheduling) cross-device remote control - is far more sophisticated from a technical depth perspective than OpenClaw. Computer Use has reached human-level performance on the OSWorld benchmark (72.5% vs human 72.4%). However, it lacks two things.

In the first quarter of 2026, AI Agent came of age. IM made Agent wait for you in the interface you're most familiar with. 7×24 allowed it to wake up and patrol on its own without waiting for you to speak. Combining the two, Agent no longer waited for you to speak, it took the initiative to come to you. OpenClaw didn't bother explaining to users what context windows or retrieval enhancement were, instead directly saying in plain language - "I'll be online at all times, I'll remember what you say, and I'll get things done on my own." Focusing on the efficacy before explaining the principles, this approach directly broke through the technical barriers. 22% of employees started secretly using OpenClaw without the IT department even knowing.

In the first quarter of 2026, AI Agent came of age. Accessibility trumped capability. Its technical depth may not have been as impressive as Cowork's OpenClaw, but it won over users by appearing in the right interface, at the right time, and in the right posture.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. Five forks emerged simultaneously, and OpenClaw is not the only one.

In the first quarter of 2026, AI Agent completed its coming-of-age. Among the automated products released concurrently with OpenClaw, we can see five different routes.

In the first quarter of 2026, AI Agent completed its coming-of-age. This was made possible by two conditions being met simultaneously.

In the first quarter of 2026, AI Agent came of age. First, the model finally crossed the "sustainable execution" threshold. The current model still makes mistakes, but at least it can hold on for dozens of cycles without suddenly forgetting what it's doing halfway through. This difference is crucial - local errors can be corrected by the system's scaffolding, but a global collapse is incurable.

In the first quarter of 2026, AI Agent came of age. Secondly, the Harness engineering methodology has become sufficiently stable. Memory has transitioned from a black box vector database to plain text files that users can directly browse and edit, supporting Git version control. The execution environment now features a gateway, heartbeat mechanism, browser takeover, and remote node invocation.

In the first quarter of 2026, AI Agent came of age. With enhanced capabilities and scaffolding in place, long-range Agent became the industry's common choice. The Worktree architecture of Codex App enabled multiple Agents to work in parallel in the same code repository, with 5 parallel Worktrees reducing a 42-minute task to 14 minutes and zero merge conflicts. The execution span of Agent officially transitioned from minutes to days.

In Q1 2026, AI agents came of age. After coding, a second must-have use case has emerged.

In the first quarter of 2026, AI Agent came of age. With OpenClaw and the Skill market, Agent transformed from a developer tool to a versatile work assistant, capable of handling tasks such as research, monitoring, content generation, and customer service. The 13,700 Skills covered a wide range of scenarios, far exceeding what could be achieved through coding alone.

In the first quarter of 2026, AI Agent came of age. The pattern is clear: as long as there is a Skill that standardizes the domain know-how, any long-term high-cognitive task is a piece of cake for Agent. However, short-term low-cognitive operations are not - ordering a cup of milk tea, for example, can be done in 30 seconds with a mobile phone, while Agent is actually slower.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. The walled garden could not stop Agent.

In the first quarter of 2026, AI Agent came of age. Domestically, nine major manufacturers each tied their own IM to compete for entry points. The battle was over "which app AI Agent should be integrated with", and the stakes were the right to control ecosystem entry points. However, by the end of Q1, this barrier began to loosen, with QClaw supporting Feishu and DingTalk, and OpenClaw-CN having built-in adapters for the five major IM platforms.

In Q1 2026, AI Agents came of age. Because when an Agent needs to simultaneously handle customer messages on WeChat, team collaboration on Feishu, and approval workflows on DingTalk, it cannot survive confined to a single IM silo. The more indispensable Agents become, the greater the cross-platform pressure—and the harder it gets for walled gardens to hold. The competition used to be about "who sits closest to the user." Going forward, it's about "who can make Agents work seamlessly everywhere."

In the first quarter of 2026, AI Agent came of age. The approaches of Silicon Valley and China are vastly different. Silicon Valley's battle revolves around "model providers vs intermediaries" - in mid-February, Google suddenly and massively banned users who accessed Gemini via OpenClaw, without any prior warning, shutting down hundreds of paid accounts overnight. The surface reason was "malicious use resulting in computational load far exceeding expectations", but in reality, it was OpenClaw's heartbeat mechanism, which checked the complete context with tens of thousands of tokens every 30 minutes, resulting in a single Ultra subscription user's actual consumption being equivalent to $1,000-3,600 in API prices, far exceeding the $250 monthly fee. This is a direct blow to the subscription-based business model.

In the first quarter of 2026, AI Agent came of age. Anthropic directly labeled this behavior as "Token Arbitrage", requiring users to access through API keys (priced at 5-10 times the subscription fee), and ultimately directly banned Openclaw's subscription-based entrance in early April. OpenAI, on the other hand, took the opposite approach, acquiring OpenClaw's founder and then adding it to the whitelist.

In Q1 2026, AI Agent came of age. Put simply, when an open-source middleware layer lets users bypass official pricing to access model capabilities, platforms are forced to choose between blocking it or absorbing it.

In the first quarter of 2026, AI Agent came of age. The same OpenClaw that sparked a pricing debate in Silicon Valley ignited a battle for market entry in China.

In the first quarter of 2026, AI Agent completed its coming-of-age. The first wave of substitution fell on outsourcing services.

In the first quarter of 2026, AI Agent completed its coming of age. Julien Bek of Sequoia calculated that for every $1 spent by enterprises on software, they spend $6 on services, including accounting, law, IT hosting, recruitment, and insurance brokerage. Agent's pricing unit is shifting from per seat and per feature to per workflow and per outcome, but the truly lucrative opportunity lies not in replacing humans, but in replacing outsourcing contracts.

In the first quarter of 2026, AI Agent came of age. Think about it, when a task has been outsourced, it means the company has already accepted external execution, has a existing budget line, and is buying results. Replacing an outsourcing contract is equivalent to switching suppliers, whereas replacing an internal employee is equivalent to an organizational adjustment, with the former having significantly less resistance.

In Q1 2026, AI agents had their coming-of-age moment. That explains why vertical autopilot agents in legal (Harvey), medical approvals (Anterior), and insurance (WithCoverage) are scaling far faster than general-purpose agents—they're not targeting the political minefield of "AI replacing people," but rather the natural commercial zone of "AI replacing outsourcing." The same pattern shows up with OpenClaw: the first batch of tasks individual users delegate to agents are exactly the ones they used to pay virtual assistants for—monitoring, research, social media management. The services market, six times the size of software, is where autopilot truly finds its footing.

In the first quarter of 2026, AI Agent completed its coming of age, with its technological capabilities surpassing organizational interfaces.

In the first quarter of 2026, AI Agent came of age. Block, the parent company of Square and Cash App, showcased an extreme form, where the company was directly reconstructed into a four-layer intelligent architecture, with middle management eliminated, and product roadmaps automatically generated by failure signals from the intelligent layer, no longer requiring product managers to decide on features. However, Block's premise is based on the high-frequency structured data of a two-sided transaction platform, which most companies do not possess.

the a of a minimal " placement only minimal " placement only placement incremental " placement only common data of a minimal " placement only common data of a incremental only common data of a. " ( that a of a common data of a. " "( " "( " "( " " " " " " " " " " " " " " " " " " " the a of a " " " " " " " " " " " " " " " " " " " " ". " " " " " " " " " " " " "," " " " " " " " " " " " " " " " " " " " " " " " " " " a " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " ". " " " " " " ": " " " " " " " " " " " " " " " " " " ". " " " " " " " of a " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " "

In the first quarter of 2026, AI Agent came of age. Model capabilities are a function of computing power and capital, and throwing money at them can boost their performance. However, these three factors are also a function of time and institutions, and they will not automatically keep pace just because the model becomes stronger. From copilot to autopilot, there is likely a transitional form that is being severely underestimated - AI autonomously executing tasks within clear boundaries, with those boundaries defined by business rules rather than prompts, and automatically escalating exceptions to humans. This is closer to making money than pure autopilot and more valuable than pure copilot.

In the first quarter of 2026, AI Agent came of age. Being able to act independently is the starting point for flywheels. However, as soon as Agent got on the road, its fatal weaknesses were exposed. OpenClaw's 512 security vulnerabilities, 341 malicious Skills, and bills running into hundreds or even thousands of dollars were obvious pitfalls.

In Q1 2026, AI agents came of age. Their newfound independence created new problems, and those problems in turn gave rise to the flywheel's second force.

In the first quarter of 2026, AI Agent came of age. 02 Constraint engineering allowed Agent to learn discipline.

In the first quarter of 2026, AI Agent came of age. The number one problem exposed after Agent became capable of independent action was that it didn't follow the rules. With a memory like a goldfish, it would declare a task complete after just three steps, give itself high scores, but ultimately fail to deliver end-to-end results.

In Q1 2026, AI Agent came of age. The quarter produced a solution in just 15 weeks. From Anthropic's first Harness blog post on December 5, 2025, to LangChain's generalized definition on March 10, Harness Engineering (constraint engineering) achieved industry consensus.

In the first quarter of 2026, AI Agent completed its coming-of-age. This speed in itself shows how anxious everyone is, as Agent has already been launched, but the rules have not yet been established.

In the first quarter of 2026, AI Agent completed its coming of age, transitioning from "what to look at" to "how to continue doing things right".

In the first quarter of 2026, AI Agent completed its coming of age. Context Engineering is in charge of the information layer, namely "what the model should see". Harness Engineering is in charge of the structural layer, namely "how the model continues to do things correctly over dozens of rounds".

In the first quarter of 2026, AI agents came of age. That maturation was the real paradigm leap of Q1.

In the first quarter of 2026, AI Agent came of age. Claude Code made 135,000 to 326,000 public commits on GitHub every day, accounting for 4% of global public commits, and is expected to reach 20% by the end of the year. The agent has become so deeply embedded in code repositories that it lacks a dedicated set of disciplinary constraints, and it's only a matter of time before problems arise.

In the first quarter of 2026, AI Agent marked its coming of age. Tencent Senior Executive Vice President Tong Dawei set the tone for the Chinese version at the Shanghai Summit on March 27: "The implementation of AI is not just an algorithmic problem, but also an engineering problem. As the capability gap between mainstream large models narrows, the core of enterprise competition is no longer the strength or weakness of the model itself, but the ability to unleash the value of the model through engineering means."

In the first quarter of 2026, AI Agent came of age. Other major tech companies have also piled in. ByteDance's DeerFlow 2.0 explicitly uses the term "Super Agent Harness" in its GitHub description—likely the first time a Chinese open-source project has adopted that phrase in its product positioning. Within a month, its GitHub stars surged from 22K to 52K, a testament to the market's voracious appetite for harness solutions.

AI agents came of age in the first quarter of 2026. Harness is a three-layer shell.

In the first quarter of 2026, AI Agent completed its coming-of-age. To put Harness into perspective, we can imagine Agent as a car. The model is the engine, and the Prompt is the steering wheel. However, having just an engine and a steering wheel does not make a car - you also need to install a transmission, dashboard, and brakes. Harness refers to these components, and currently, in engineering practice, there are three layers in total.

In the first quarter of 2026, AI Agent completed its coming-of-age rite. The first layer is process control, designed to rein in disobedience. Agents have memories like goldfish, declare tasks complete after just three steps, and remain oblivious when the environment throws bugs. The remedy: externalize state (via AGENTS.md / progress files), decompose tasks, and enforce step-by-step execution.

In the first quarter of 2026, AI Agent completed its second-layer development, concurrent scheduling, which specializes in preventing group "phishing" attacks. When 100 Agents run simultaneously, they can easily avoid risks collectively and focus on making the simplest minor modifications, while neglecting the truly difficult problems. The solution is to implement a multi-layer Agent structure, separate roles (Planner for planning → Generator for generation → Evaluator for evaluation), and anti-phishing mechanisms.

In Q1 2026, AI agents came of age. The third layer is verification and error correction, designed to counter unwarranted self-assurance. When AI grades its own work, Anthropic calls this "self-deception." The remedy: independent evaluators, sandbox isolation, and Git transaction boundaries—branch as sandbox, PR as approval, merge as the only true commit.

In the first quarter of 2026, AI Agent came of age. CLI, Skill, and externalized memory formats, which are often mentioned, also became trends in the development field in the first quarter. In engineering practice, it's not entirely accurate to equate this with Harness, but rather with Agentic Infra (Agent infrastructure). Harness is only concerned with "how to drive steadily," while Infra is concerned with "road conditions and gas stations" that can accelerate and expedite progress.

In the first quarter of 2026, AI Agent completed its coming-of-age. Every layer of the shell was squeezed out by bugs.

In the first quarter of 2026, AI Agent came of age. Harness did not emerge out of thin air, but rather from a series of practical engineering problems. As people wanted models to perform longer-term tasks, bug after bug drove its development.

In the first quarter of 2026, AI Agent completed its coming of age. The first version's origin. Users would give Agent a large demand all at once, and it would try to complete it all in one go, collapsing at the 30th step. Anthropic came up with a simple solution, like a relay race - "initialize Agent" sets up the environment, writes a handover list (claude-progress.txt), and then exits. "Coding Agent" reads the handover list each time it comes on, figures out where the previous one left off, only does one function, updates the list after completion, and then exits. The key point is that Agents do not share conversation history, only transmitting information through files. This is because the conversation history is completely overwhelmed by noise from the previous nine rounds by the tenth round.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. Cursor discovered that Agent, under a flat structure, was extremely risk-averse, preferring to make meaningless minor modifications rather than tackling difficult issues, resulting in the entire system being idle. Anthropic introduced a "client-server" architecture - Planner wrote specifications, Generator implemented each function according to the specifications, and before starting each function, a Sprint contract was written. Evaluator tested functions using a real browser (not just code), scoring them based on four dimensions: product depth, functionality, visual design, and code quality. If the score was not up to standard, the sprint failed and rework was required. An interesting finding was that it was much easier to adjust the "scorer" to be stricter than to teach the "coder" to be self-critical.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony, revealing the origin of its third layer. Agent itself ran the tests it wrote and claimed "no bugs," but the end-to-end process was fundamentally inoperable. Anthropic refers to this as "self-deception" - similar to letting students grade their own essays, the score will never be low. It must have an independent Evaluator and sandbox isolation.

In the first quarter of 2026, AI Agent completed its coming of age. Mitchell Hashimoto's AGENTS.md in the open-source project Ghostty is not a design document, but rather a record of accidents - whenever Agent modifies a file it shouldn't, a new rule is added, such as "do not modify the vendor/ directory". If Agent uses an outdated interface, a new rule is added, such as "use v2 API instead of v1". If Agent writes nonsense in commit messages, a format specification is added. He found that on normal working days, the background Agent could only run effectively for 10-20% of the time, and at first, "it took longer than doing it manually". However, once it reached a turning point, and the rules accumulated to a certain density, the error rate of Agent significantly decreased.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. OpenAI discovered a more chronic disease without a dedicated maintenance mechanism, where the repositories used by Agent would significantly deteriorate after about 2-3 months. It's like having ten interns come in to work every day, leaving behind a pile of "temporary solutions" when they leave, and after three months, no one can tell which code is seriously written and which is makeshift. They set up three mechanisms - architecture constraint declarations (clearly writing out the project's framework, model, and naming specifications in AGENTS.md), Doc gardening (regularly cleaning up Agent's outdated comments and redundant documents), and Anti-slop routine (anti-degradation inspection, cleaning up Agent's accumulated inconsistent styles and duplicate code). It's not enough to just do it once, it needs to be continuously run.

In the first quarter of 2026, AI Agent completed its coming-of-age. Upgrading the shell is more cost-effective than upgrading the model, but not cheap.

In the first quarter of 2026, AI Agent came of age. LangChain conducted an experiment where the same model was used with a different Harness, and the pass rate for Terminal Bench 2.0 jumped from 52.8% to 66.5% without any change to the model's weights, with its ranking soaring from outside the top 30 to the top 5.

Q1 2026 marked the coming of age for AI agents. That's the Harness effect.

In the first quarter of 2026, AI agents completed their rite of passage. But that milestone came at a steep price.

In the first quarter of 2026, AI Agent came of age. According to Anthropic's cost data, Solo Agent can barely run the same 2D game, spending $9 and 20 minutes, but the resulting product is mainly functionally damaged and unplayable. With a complete Harness, it costs $200 and 6 hours, producing a finished product with complete functionality, exquisite visuals, and normal gameplay.

In the first quarter of 2026, AI Agent completed its coming of age. A 20-fold increase in cost did not result in a slight improvement, but rather a difference between being usable and unusable, a matter of life and death.

In the first quarter of 2026, AI Agent came of age. From this point on, Harness is currently the most cost-effective capability amplifier, but it is not cheap, and it is not on the same level as maintaining a continuously running, well-behaved Agent and occasionally inquiring chat assistant.

In the first quarter of 2026, AI Agent completed its coming-of-age. However, bills ranging from several hundred to several thousand dollars have directly discouraged many users after just a few weeks of experience. Cost remains the biggest obstacle to popularization.

In the first quarter of 2026, AI agents came of age. Harness is a temporary moat, but the compensation landscape is shifting.

In the first quarter of 2026, AI Agent came of age. While the entire industry was busy building bricks for Harness, Anthropic was already breaking down the walls it had built itself.

In the first quarter of 2026, AI Agent came of age. After the release of Opus 4.6, they removed Context Reset, as the model's context management capability had become strong enough to render context resetting unnecessary. They also removed Sprint Contract, as the new model was able to control the pace on its own, eliminating the need for a acceptance contract to be signed at the start of each round. The Evaluator was also changed from evaluating each round to only conducting QA in the final round.

In the first quarter of 2026, AI Agent came of age. In Anthropic's own words, "Every component of Harness encodes an assumption about what the model can't do. When the assumption no longer holds, the component is due for a change."

In the first quarter of 2026, AI Agent completed its coming-of-age. Being able to dismantle and explain what was initially set up effectively shows that they have always been clear about what they were compensating for.

In the first quarter of 2026, AI Agent completed its coming-of-age. The difficulty lies not in dismantling, but in determining when to dismantle. If dismantled too early, the model may not be able to support itself and the system will collapse; if dismantled too late, the shell will conceal the model's true capabilities.

In the first quarter of 2026, AI Agent came of age. As the model evolved, you thought the shell was helping, but in fact, it was getting in the way.

In the first quarter of 2026, AI agents came of age. The path to simplicity runs through complexity. But so far, only Anthropic has completed the full cycle—from adding to stripping away.

AI agents came of age in the first quarter of 2026. Beyond the harness lies a bigger shell.

In the first quarter of 2026, AI Agent came of age. Five key consensus points were widely discussed in Q1: Markdown serving as a state carrier, Git as a transaction boundary, the resurgence of CLI (Agent requires a structured text interface, with a single git diff --stat command providing an overview of all changes), real-time testing (Agent modifies a file and immediately runs tests, with test results becoming instant feedback signals at each step), and Skill as a knowledge encapsulation, which actually does not entirely belong to Harness.

In the first quarter of 2026, AI Agent completed its coming of age. The larger framework is called Agentic Infra (Agent Infrastructure), which is divided into five layers: the Context layer (what the Agent can remember), the Tool Interface layer (what the Agent can do), the Harness layer (how the Agent is managed), the Knowledge layer (how the Agent knows what to do), and the Economic layer (how much it costs to run the Agent).

In the first quarter of 2026, AI Agent completed its coming-of-age. Harness is the most core layer, but it's far from the only one. Currently, the industry's attention is focused on the Harness layer because it directly determines whether the Agent can be used. However, the next stage of the battle may be fought on other layers - the Context layer's memory quality (memory retention across days and sessions is still unreliable), the tool interface layer's execution environment resources (Anthropic found that simply relaxing resource constraints can increase the success rate by 6 percentage points), the knowledge layer's Skill triggering mechanism (Vercel's evaluation shows that in 56% of cases, the Agent won't proactively use its existing Skills), and the economic layer's cost control (model routing, budget allocation, and parallelization strategies are still largely based on intuition).

In the first quarter of 2026, AI Agent came of age. However, there is another issue that everyone is aware of but has no good answer to - organizational governance. "Who approved the code written by Agent?" In most companies, there is no standard answer. Audit logs, decision tracing, and code ownership are all concerns. A three-person startup team may be able to get by with approval checkpoints at the workflow level, but in a company with 500 people, without an independent governance framework, it is impossible to rely solely on technical means.

Q1 2026: AI Agents Come of Age This is the part of the Harness architecture that isn't directly responsible for the core task, yet is the crucial element that makes agents truly dependable.

In the first quarter of 2026, AI Agent reached its coming of age. Agent failures can finally be diagnosed.

In the first quarter of 2026, AI agents completed their rite of passage. A three-layer shell and five-layer infrastructure gave the industry its first diagnostic framework. When an agent fails, you can no longer just dismiss it with "the model isn't good enough."

In the first quarter of 2026, AI Agent completed its coming of age. Was it due to inadequate process control, uncontrolled concurrent scheduling, missing verification links, insufficient execution environment resources, untriggered Skills, or fundamentally unprofitable costs?

" as of top of as of top of a as of top of a top of top of a. " ( as of a top of a. " "s of top of a. " " "s of a top of " "s of a top of " " "s of a. " " "s of that top of " " "s of first top of " " "s of common " " "s of common top of " " "s the a of top of a. " " "s a of top of a. " " "s a of a top " " "s only minors of " " "s only common top of a. " " "s only minors of " " "s only common only top of a. " " "s only minors of " " "s only " " "s only common only. " " "s only minors of " " "s only common only top of a. " " "s only minor of a. " " of a. " " "s only minor of " " "s only common only. " " "s only minors of a. " " "s only minores of a. " " "s only minors

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. Constraint engineering gave Agent discipline, and the flywheel can now turn into the next cycle.

In the first quarter of 2026, AI Agent came of age. A disciplined Agent finally possessed an ability that was previously impossible, continuously improving itself in long-term cycles, rather than collapsing at the tenth step.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony, which is the premise of the third force of the flywheel.

In the first quarter of 2026, AI agents completed their rite of passage. 03 Recursive R&D: Agents Begin to Strengthen Themselves

In the first quarter of 2026, AI Agent came of age. The previous two chapters discussed how Agent established itself as a product and system. This chapter explores the scenario in which Agent, having gained discipline, first broke through its role as an "executor" and began to improve its execution methods.

In the first quarter of 2026, AI Agent came of age. The answer is R&D. Because R&D is naturally verifiable (passing tests means passing), reversible (Git can be reverted with one click), and readable and writable (code itself is plain text that machines can directly operate on).

In the first quarter of 2026, AI Agent completed its coming of age. Three conditions came together, enabling Agent to enter a complete cycle of "execution → verification → problem discovery → modification → re-execution".

In the first quarter of 2026, AI Agent came of age. There are three paths here, all with dazzling result data. AlphaEvolve recycled 0.7% of global computing power, Minimax M2.7 saw a 30% improvement in internal evaluation after over 100 rounds of autonomous iteration, and Karpathy's Autoresearch ran 50 experiments in one night.

In the first quarter of 2026, AI Agent came of age. These three open-sourced practices demonstrate that recursive development is now generating tangible value. Meanwhile, major model manufacturers that have not open-sourced their technology have also acknowledged this development in various interviews.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. Exploration, optimization, and engineering workflows are three completely different recursions.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. Now, let's take a closer look at these three methods.

In the first quarter of 2026, AI Agent came of age. Exploratory in nature, AlphaEvolve doesn't tune parameters—it searches for algorithms that humans have never seen. It's an evolutionary system built on Gemini Flash, which handles breadth by rapidly generating large numbers of variants, and Gemini Pro, which handles depth by meticulously refining the best solutions. The entire process produces human-readable code, not a black box—understandable, tunable, and ready for deployment. The data center scheduling algorithm it discovered has been running in Google's production environment for a year, continuously recovering 0.7% of global compute capacity—worth billions of dollars. It proposed optimizations for critical TPU circuits, speeding up a key computation kernel in the Gemini architecture by 23% and optimizing FlashAttention's low-level instructions by 32.5%. Of more than 50 open problems in mathematics, it improved the best-known solution for 20% of them, including improving upon the matrix multiplication algorithm Strassen introduced in 1969. The recursive value here is that it could change the direction of an entire discipline.

In Q1 2026, the AI agent came of age. Optimization-type agents, Autoresearch and M2.7. With a known objective function, the task is simply iterative refinement toward an optimum. Karpathy distilled the core loop into 630 lines of Python—three files (train.py can be modified, prepare.py is fixed, program.md holds human-written instructions for the agent), plus a "ratchet" rule (keep only results better than the last run; never regress). Roughly 12 experiments per hour, 80–100 overnight. 23K GitHub stars in three days, 35K in three weeks. The pattern has since moved beyond ML—LangChain founder Harrison Chase applied the same three-file architecture to optimize LangChain's agent itself, and others have used it for database query optimization and customer-support ticket routing. Any optimization problem where quality can be quantified fits the template.

In the first quarter of 2026, AI Agent came of age with M2.7 and further advancements. MiniMax enabled the model to act as a "research-type Agent," taking charge of improving its own reinforcement learning training process. It autonomously constructed dozens of complex Skills, updated its memory system, and continuously optimized the entire Harness architecture. The model's memory is divided into three layers: after each iteration, it writes a "short-term note" (recording what it did and the results), conducts "self-criticism" (identifying what went wrong and how to improve next time), and then reviews all historical data before deciding on the next direction. This process is similar to a researcher writing an experimental log and reflecting on it every day, rather than starting from scratch. After 100+ rounds of autonomous iteration, internal evaluation metrics improved by 30%, and the SWE-Pro score reached 56.22%, matching GPT-5.3-Codex. In 22 ML competitions, the model achieved 9 gold, 5 silver, and 1 bronze medals after three rounds of self-optimization. The API price is only 8% of Claude 4.5 Sonnet's. This recursive approach accelerates progress along existing

In the first quarter of 2026, AI Agent came of age. Engineering models such as Codex participated in OpenAI's internal research and development, while Claude wrote Anthropic's own code. Dario Amodei confirmed that over 90% of new code was written by AI itself. This recursive value is most straightforward, releasing human resources and accelerating iteration.

In Q1 2026, AI agents reached their coming-of-age moment. The three types of closed loops are fundamentally different in nature—but agents are no longer just "doing the work." They are now improving the very methods by which they work.

In the first quarter of 2026, AI Agent came of age. The speed of the human brain has become the system's bottleneck.

In the first quarter of 2026, AI Agent marked a significant milestone. At Autoresearch, humans have completely exited the execution loop, only needing to set up three files for the Agent to run on its own. However, the entire process remains human-in-the-loop, with two tasks that AI cannot automate: defining objectives ("which metrics to optimize") and determining boundaries ("which directions to avoid").

In Q1 2026, AI agents reached their coming-of-age moment. When an agent runs 50 rounds in a single night and 500 rounds in a single day, humans can no longer keep pace with target-setting based on gut instinct. The bottleneck for self-evolution—which still requires a human in the loop—has shifted from "not enough hands" to "not enough brain speed."

In the first quarter of 2026, AI Agents reached their coming-of-age moment. At the Zhongguancun Forum, Yang Zhilin of Moonshot AI said, "AI will define the most appropriate reward function for a given environment, and even explore new network architectures." Xiaomi's Luo Fuli added, "By layering a verifiable constraint condition and a Loop onto the existing Agent framework, the model never stops." "Our team is already using this approach for research tasks, with efficiency gains of nearly tenfold."

In the first quarter of 2026, AI Agent came of age. It shortened its judgment of the time required for AI to achieve complete self-evolution from "three to five years" to "one to two years".

In the first quarter of 2026, AI Agent completed its coming of age. Two parties are trying to break through the same bottleneck. The ultimate question is who has the agenda-setting power. Autoresearch is a "faster experimental assistant", setting human-designed goals, with Agent executing. "Self-evolution" is AI itself deciding the research agenda, setting its own goals, designing experiments, running, evaluating, and adjusting direction. The gap lies not in technical capabilities, but in who holds the steering wheel.

In the first quarter of 2026, AI Agent completed its coming-of-age. The acceleration of recursive R&D is exponential, not linear.

In the first quarter of 2026, AI Agent came of age. Mimimax M2.7 optimized its own toolchain with its own output, and the optimized toolchain enabled the next round to be more efficient. The more efficient next round produced an even better toolchain. This demonstrates that the R&D speed driven by AI self-evolution is compound, rather than linear.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. However, compound interest has a fatal premise: each round of improvement must be "real improvement" and cannot be fabricated. If the evaluation pipeline itself is biased, the compound interest will become "running faster and faster in the wrong direction". This is the same problem as Goodhart's Law, where when a metric becomes a target, it ceases to be a good metric.

In the first quarter of 2026, AI Agent came of age. The scheduling algorithm discovered by AlphaEvolve took a whole year of running in a production environment to be confirmed effective, but most teams couldn't wait that long. There is a structural mismatch between short-term evaluation metrics and long-term true value.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. Evaluating the authenticity of the pipeline determines whether compound growth can be sustained.

In the first quarter of 2026, AI agents came of age. The next battleground is self-evolving infrastructure.

In the first quarter of 2026, AI Agent came of age. Over the past decade, the competition was about training infrastructure, with companies vying to have the largest computing power clusters, the fastest data pipelines, and the most stable training frameworks. However, the focus of competition has now shifted to something new, called self-evolving infrastructure.

In the first quarter of 2026, AI Agent completed its coming of age. Based on current practices, it mainly consists of five components: the separation of variable assets and immutable infrastructure (clearly defining what Agent can and cannot modify), evaluation pipelines (quickly and accurately determining "this round is better than the last"), memory and selection mechanisms (remembering good experiences and discarding bad ones), execution environments (sandbox isolation with abundant resources, with Anthropic finding that simply relaxing resource constraints improved performance by 6 percentage points), and dynamic tools and skills (M2.7's Agent created dozens of auxiliary tools for itself to run reinforcement learning experiments).

In the first quarter of 2026, AI Agent came of age. AlphaEvolve can recycle 0.7% of global computing power, not just because the Gemini model is strong, but also because Google has the world's best evaluation pool and parallel execution environment. As model capabilities begin to converge, the gap in self-evolving Infra is where the true difference lies.

In the first quarter of 2026, AI Agent completed its coming of age. Agent learned to grow on its own, with the flywheel turning into its third cycle. However, recursive development exposed a new bottleneck - Agent accumulated experience from scratch in each cycle, while decades of industry know-how possessed by humans was left untapped.

In the first quarter of 2026, AI Agent came of age. If there is a format that allows Agent to directly inherit the experience of its predecessors, the starting point of recursion will no longer be zero, but the endpoint of its predecessors.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. Q1 happened to appear in this format.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony, the 04 Skill ecosystem, where knowledge is no longer tied to individuals.

In the first quarter of 2026, AI Agent came of age. Opus 4.6 can write code in any language, but it is unaware of your team's coding standards, unclear about your industry's approval processes, and even more unaware of where the technical debt of your project is buried.

In the first quarter of 2026, AI Agent came of age. "There's a hidden rate limit for this API in high-concurrency scenarios," "The migration tool for this framework has a bug prior to v3.2, and you must manually modify a configuration first," and "Our team never uses ORM's cascade delete because it caused a major accident three years ago." All of these are know-how gained by seasoned engineers through trial and error, and are not included in the training data, nor are they suitable for hardcoding into product logic.

In the first quarter of 2026, AI Agent came of age. In Q1, these experiences were packaged for the first time in a format that can be packaged, distributed, and reused infinitely, called Skill.

In the first quarter of 2026, AI Agent completed its coming-of-age. Skill fills the gap of experience, not the gap of technology.

In the first quarter of 2026, AI Agent completed its coming of age. A Skill is neither a document nor code, but a structured knowledge package that includes trigger conditions (the scenario in which it is used), standard operating procedures (step-by-step instructions), executable scripts (directly runnable tools), and reference materials (background knowledge). What it does is simple: it converts the "things in an old employee's brain" into a format that Agent can read and execute.

In the first quarter of 2026, AI Agent completed its coming-of-age. The Prompt solution addresses the issue of "how to say it more clearly" and has a sense of immediacy but is not reusable. Workflow refers to a deterministic process orchestration that is stable but rigid. Skill falls somewhere in between - more stable than Prompt (structured and version-controlled), more flexible than Workflow (the model can apply it flexibly according to the current situation), and lighter than retraining the model (modifying a Markdown file vs retraining a large model with tens of billions of parameters).

In the first quarter of 2026, AI Agent came of age. It's not that it's smarter, but rather more savvy.

In the first quarter of 2026, AI Agent completed its coming-of-age, with the ability to write once and reuse infinitely.

In the first quarter of 2026, AI Agent came of age. In the past, the transfer of experience in various fields relied on masters teaching apprentices, writing documents, and providing training, which was slow, unable to be scaled, and heavily dependent on individuals.

In the first quarter of 2026, AI Agent came of age. Now, a seasoned engineer can write a TDD Skill in two hours, and thousands of Agent instances across the company can load it simultaneously, mastering it instantly. The domain experience that used to take a junior engineer two years to accumulate can now be packaged into a file and distributed. Knowledge is no longer tied to individuals, but to the structure.

In the first quarter of 2026, AI Agent came of age. Users of the popular Skill Superpowers framework (143,000+ installations, 93K stars on GitHub) say, "Skills are not suggestions, but structured decision trees. They give Claude discipline." "I've become lazy with TDD, now skills remember for me." What skills achieve is not making the Agent smarter, but making it more reliable.

In the first quarter of 2026, AI Agent completed its coming of age. By looking at the Brainstorming Skill in Superpowers, it's possible to understand the difference between a skill and ordinary prompt words. The Agent has a tendency to directly start writing code upon receiving a vague requirement, only to discover halfway through that it has misunderstood, and then has to restart. The Brainstorming Skill inserts a hard barrier between the Agent and the code, strictly prohibiting any code from being written before the design has been approved by the user. It defines a 9-step execution checklist, ranging from exploring the project context to proposing 2-3 solutions (including weighing and recommendations), and finally to user review.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. The prompt provided was "principle" (which the Agent could choose to ignore), and the skill provided was "access control" (without which it could not proceed to the next step).

In the first quarter of 2026, AI Agent marked a major milestone. Skills can now be linked together, with Brainstorming automatically triggering the generation of writing-plans, which in turn call executing-plans, utilizing test-driven-development during the execution process.

In the first quarter of 2026, AI Agent came of age. This makes Skills not isolated capability packages, but standardized modules that can form a complete workflow.

In the first quarter of 2026, AI agents came of age. Openness and security are mutually exclusive.

In the first quarter of 2026, AI Agent completed its coming-of-age. In terms of its ecosystem, three routes are running simultaneously.

In the first quarter of 2026, AI Agent came of age. ClawHub took an open-community approach, achieving rapid growth, accumulating over 13,700 Skills in just half a year, with a single Skill reaching as many as 180,000 installations. Popular Skills naturally formed a hierarchy, including a survival layer (Web Browsing with 180,000 installations), an efficiency layer (Telegram Bot with 145,000 installations), and an advanced layer (Capability Evolver with 35,000 installations). In the advanced layer, Agent automatically identifies repetitive patterns and creates new Skills, equivalent to giving AI a "self-evolution" button.

In the first quarter of 2026, AI Agent came of age. ClawHub also achieved cross-school compatibility, supporting direct import of plugin packages from the three major platforms Claude, Codex, and Cursor, with automatic mapping and execution, and over 4,000 cross-platform skills are interoperable, firmly establishing its position as the "omnibus intermediary layer".

In the first quarter of 2026, AI Agent completed its coming of age. However, the cost of openness has also arrived. According to relevant research, in a screening of over 1,200 skills, 341 malicious skills (accounting for 11.3% of the market) were discovered, with 36% containing prompt word injection. VirusTotal directly defined this as "the AI version of npm poisoning".

In the first quarter of 2026, AI Agent came of age. What's even more frightening is the supply chain contamination incident at the beginning of the year, in which attackers triggered the AI diversion robot to execute malicious code simply by crafting a post title, poisoning the cache, stealing tokens, and ultimately forcibly installing backdoors on the machines of several thousand developers.

In the first quarter of 2026, AI Agent came of age. The MCP ecosystem suffered 30 CVEs within 60 days, with 82% having path traversal vulnerabilities and 38% lacking any authentication.

In the first quarter of 2026, AI Agent came of age. When a Skill invokes external tools via MCP, two layers of risk are overlaid and amplified.

In the first quarter of 2026, AI Agent came of age. In response, Chinese manufacturers have placed considerable emphasis on security. Tencent's SkillHub takes a platform audit approach, which is secure but has limited openness. Dui Zi and DeerFlow have taken an open-source, controllable approach. Skills are Markdown files that can be version-controlled in Git, and loaded on demand without occupying the context window.

In the first quarter of 2026, AI Agent came of age. The format of Skill has been established, and now the competition in the ecosystem is over who will distribute and how to distribute. The 341 malicious Skills indicate that this ecosystem is still very immature. However, "immature" and "non-viable" are two different things.

In the first quarter of 2026, AI Agent marked a significant milestone. 56% of Agents would not proactively check their Skills.

In the first quarter of 2026, AI Agent marked its coming of age. It has acquired skills, but its integration with the system still lacks maturity.

In the first quarter of 2026, AI Agent came of age. Vercel conducted a highly precise evaluation using the new Next.js 16 API, which was not included in the model's training data. Without providing any information, the pass rate was 53%. After providing an AGENTS.md index file (directly inserted into the system prompt), the pass rate soared to 100%. However, when given a Skill (placed on a shelf for it to find on its own), the pass rate remained at 53%, similar to when no information was provided.

In the first quarter of 2026, AI Agent came of age. However, in 56% of cases, Agent was not even aware that it needed to search for information. No matter how many good Skills were available in the market, if Agent itself did not know to look for them, it was equivalent to not having them at all.

In the first quarter of 2026, AI Agent completed its coming of age. DeerFlow's solution is to explicitly load Skills when decomposing tasks at the orchestration layer - rather than relying on the Agent to search for them itself, the system decides for it during the planning phase. This actually pushes the problem back to the Harness workflow layer.

In the first quarter of 2026, AI Agent came of age. The value of Skill will continue to be severely underestimated until its triggering mechanism is mature.

In the first quarter of 2026, AI Agent completed its coming-of-age. What Skill is shaking is not the interface layer, but the process layer itself.

In the first quarter of 2026, AI Agent came of age. The core value of SaaS lies in solidifying domain workflows into software. Essentially, a CRM is the software version of "customer management know-how".

In the first quarter of 2026, AI Agent came of age. Skill does the same thing, but at a cost several orders of magnitude lower (writing Markdown vs developing a SaaS suite), with iteration several orders of magnitude faster (changing one line vs releasing a version), and distribution several orders of magnitude faster (Agent loading directly vs user registration, learning, and data migration).

In the first quarter of 2026, AI Agent came of age. MCP had once shaken up SaaS, but it only affected the interface layer, with the underlying processes remaining intact, and while its business model was challenged, it was not fundamentally disrupted. Earlier, SFT had also caused a stir, but it was inaccessible to ordinary people and failed to generate scalable compound benefits.

In the first quarter of 2026, AI Agent came of age. Its Skill is different, as it shakes up the process layer itself. When a Skill enables an Agent to run the entire "manage customers with Salesforce" process, users no longer need the Salesforce interface. The threshold is extremely low (writing in Markdown), and it can accumulate compound interest (over 13,700 in half a year).

In the first quarter of 2026, AI Agent came of age. With the maturation of Skills, the next to face threats after SaaS may be Apps. When an Agent can complete tasks like "ordering takeout + price comparison + reaching a discount threshold" through Skill combinations, will you still need to open Meituan?

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. AI Agent's App Store is being born.

In the first quarter of 2026, AI Agent came of age. GPT Store, which sells chatbots and simple workflows, struggled to gain traction as they were not essential products. ClawHub, on the other hand, sells "capability packages" that enable Agents to perform specific tasks, which are essential for Agents that run continuously. OpenClaw released an epic update at the end of Q1, featuring 45 new functions, 13 breaking changes, and 82 fixes. The Agent timeout period was extended from 10 minutes to 48 hours, and a pluggable sandbox backend architecture was added, along with three major search services.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony. The Skill market, in this context, has become an emerging App Store. Unlike any previous attempts, this time it has grown on the soil of Agent's continuous operation.

In the first quarter of 2026, AI Agent completed its coming-of-age. Long-range Skill will repeat the mistakes of Harness.

In the first quarter of 2026, AI Agent completed its coming of age. Short skills (such as "indenting with 2 spaces") are almost error-free. However, for long-form skills (such as writing a complete marketing plan based on the AIDA model, exceeding 2000 characters and involving over a dozen steps), the Agent often skips steps, fails to follow through, and forgets previous requirements halfway through. These failure modes are identical to those encountered when Harness was first launched.

In the first quarter of 2026, AI Agent marked a milestone. M2.7 maintained a 97% single-step follow-through rate in 40 complex Skill tests, which sounds impressive. However, relevant research shows that a 3% failure rate can accumulate in long-term tasks, resulting in a significantly lower overall success rate - with 10 steps at 97% each, the overall success rate drops to 74%. Skills require their own Harness, including step-by-step breakdown, progress tracking, intermediate verification, and rollback mechanisms. But currently, no one is working on this.

In the first quarter of 2026, AI Agent completed its coming of age. This is a known structural defect, waiting for the first batch of pioneers to be forced to come up with a solution, just like Harness's three-layer shell was back in the day.

In the first quarter of 2026, AI Agent completed its coming of age. After experiences were distilled, where will humans retreat to?

In the first quarter of 2026, AI Agent came of age. Skill distilled human experience into a format that Agent can execute directly. After a seasoned engineer wrote a Skill, all Agents in the company could use it instantly, so what would this engineer do next?

In the first quarter of 2026, AI Agent came of age. In the short term, it "moved up" to the judgment and decision-making layer. The execution layer was handed over to the Agent, while humans retreated to defining objectives, auditing quality, and handling edge cases. This is the same issue as the human brain becoming the bottleneck, as mentioned in Chapter 3.

In the first quarter of 2026, AI Agent came of age. But a more pressing issue is that in an organization, 1,000 executors may be needed, while only 10 decision-makers are required. When Skill distills all the know-how of the execution layer, the work of 1,000 executors is replaced by Agent, and they "move up" to the decision-making layer, but the decision-making layer cannot accommodate 1,000 people. This is not a transformation of jobs, but a net reduction in the total number of jobs.

In the first quarter of 2026, AI Agent came of age. Moreover, once experiences are distilled and written into Skills, the Skills no longer need you.

In the first quarter of 2026, AI Agent came of age. Humans are stepping back from evaluation, as recursive R&D is taking over judgment (Evaluator). Humans are stepping back from creation, but AlphaEvolve is discovering algorithms that humans had never thought of.

In the first quarter of 2026, AI Agent came of age. Q1 did not answer where humans should retreat to, but it did something even more ruthless - it turned this question from a "philosophical discussion" into a "reality to be faced next quarter".

a " top of a top of a top of a " top of a a " top of a top of a " top of a " top of " " " " " " " " " " " " " " " " " " " " ". " " " " " " " " " " " " " " " " " " " ". " " " " " " " " " " " " "ing " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " an " " " " " " " " " " " " " " " " " " " ". " " " " " " " of " " " " " " " " " " " " " " " " " " " " " " " " " ". " of " " " " " " " " " " " " "." " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " "iles." " " "

In the first quarter of 2026, AI Agent completed its coming-of-age. Four driving forces, one flywheel.

In the first quarter of 2026, AI Agent came of age. Productization put Agent on the road, constraint engineering taught it discipline, recursive R&D enabled it to learn self-improvement, and skill ecology allowed it to inherit experience from its predecessors. Each force was the inevitable consequence of the previous one and the necessary premise for the next.

In the first quarter of 2026, AI Agent completed its coming-of-age. However, the most notable property of the flywheel is not causality, but acceleration.

In the first quarter of 2026, AI Agent came of age. Skills make Agent stronger → Agent can handle more complex tasks → More complex tasks drive the creation of more precise constraints → More precise constraints support deeper recursion → Deeper recursion produces better skills. Each cycle accelerates the next, resulting in exponential growth, not linear.

In the first quarter of 2026, AI Agent came of age. Q1 marked the first complete rotation of the flywheel. The speed is still not fast, and there is a lot of friction between the gears. 341 malicious Skills, a 56% Skill trigger failure rate, costs running into thousands of dollars, and a lack of organizational governance.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony, but the flywheel has already started to turn.

In the first quarter of 2026, AI Agent came of age. Among the 25 evaluations, the most important one is not any specific evaluation, but the collective conclusion they point to: Agent is no longer a tool that needs to be guided by humans hand-in-hand. It is evolving into a "new colleague" that can work independently, follow rules, grow, and understand its profession.

In the first quarter of 2026, AI Agent came of age. The arrival of this new colleague has been faster than most people expected, faster than most organizations are prepared for, and faster than most discussions about "where humans will retreat to".

In the first quarter of 2026, AI Agent came of age. Q1 did not provide an answer to where humans should retreat, but it transformed the question from a "philosophical discussion" to a "reality to be faced next quarter".

In the first quarter of 2026, AI Agent came of age. The flywheel won't wait for you to make up your mind before it starts turning.

In the first quarter of 2026, AI Agent completed its coming-of-age ceremony This article comes from WeChat public account "Tencent Technology", author: Boyang, published with authorization from 36Kr.

← Back to Latest