"OTT" has had two vastly different fates in the communications industry. In the mobile internet era, it was a nightmare for operators: internet companies used the networks built by operators to offer free messaging services through WeChat, killing off text messaging services worth tens of billions of yuan. Operators invested heavily in building networks, but only earned basic traffic fees, while high-value-added businesses were taken away, leaving them trapped in a "pipeization" dilemma.
In the era of large models, computing power, networks, and localized delivery have once again become core industrial resources, and the narrative is beginning to reverse. Operators are at a crossroads: continue to be passive in the face of over-the-top (OTT) services, or pivot to become AI infrastructure providers.
There have long been two schools of thought in the industry: one advocates for relying on network dominance to strangle AI-based OTT services, while the other advocates for fully investing in the development of general large models to compete with leading AI companies. However, reality shows that controlling access is not feasible and developing general large models does not align with the resource endowments of operators. A more reasonable positioning is to become the "pipeline" for large models - but it is essential to learn from the lessons of the first generation of OTTs: in the AI era, the pipeline must not repeat the old path of charging by traffic or token, otherwise it will fall into the trap of increasing volume without increasing revenue.
Nubia's Plan to Control OTT Entrance Fails
Whenever operators negotiate with internet platforms, the phrase "relying on network layer advantages to constrain OTT manufacturers" is mentioned. This approach is not unfeasible from a technical standpoint, but rather faces three major hard constraints that are difficult to implement.
First, regulators won't allow it. As operators possess quasi-public infrastructure attributes and undertake universal service obligations, network neutrality is a governing principle. Regulators will not permit operators to leverage their underlying networks to engage in asymmetric competition with upstream and downstream entities - business demands must be subject to industry governance logic.
Second, public opinion will not allow it. Users' expectations for network experience have stabilized, and network resources are already operating at high load. If commercialized sales of differentiated acceleration and priority guarantees are accelerated, creating differences in access levels, public opinion is likely to erupt. Under the premise that there is no room for network experience, it is unrealistic to earn excess profits by relying on privileges.
Third, the model is not viable. The essence of "chokehold" is to collect protection fees, which requires controlling user entry points - however, in the mobile internet era, C-end entry points have long been in the hands of internet applications. While B-end does have QoS guarantee services, they have always been limited to customized scenarios and cannot become strategic-level businesses. Relying on network control to reap the benefits of AI is an illusion.
Round Two: General Large Models, An Unwinnable Battle
Since control is not working, another faction advocates: investing heavily in self-developed general large models, creating its own C-end AI products. This is more like an illusion of the pragmatic faction.
Telecom operators have spent years exploring computing power and cloud services, yet their efforts have consistently amounted to much ado about nothing. Building infrastructure is both their core duty and their strength—massive data centers, bandwidth, power, and local nodes are inherent advantages of these state-owned enterprises. But once they extend into large-model product refinement, ecosystem operations, and developer services, their weaknesses become apparent: in product iteration, ecosystem building, and developer relations, operators lag far behind internet giants. This gap cannot be closed by throwing more money or headcount at it—it stems from a difference in organizational DNA.
Price is the most direct evidence. A real-world example: consuming the same amount of DeepSeek-V4 tokens, using Claude costs around 26 yuan for 120 million tokens, while a certain operator charges 30 yuan for just over 20 million tokens - several times more expensive.
What's the key point? First, we need to correct a binary misconception: operators selling tokens externally has never been an either-or choice between "reselling or building their own", but rather a hybrid model running in parallel - simultaneously operating through agent reselling, deploying open-source GPUs, and joint deployment, with each of these three models driving up costs.
The agency model (connecting to manufacturers' public APIs for resale and wholesale): operators do not run models, instead obtaining wholesale prices from manufacturers such as DeepSeek, Qianwen, and Zhipu, and packaging them into their own MaaS platform to sell to customers as suites. Customer calls go through the operator's gateway, and the operator settles with upstream providers on a pay-per-use basis. Each call incurs an upstream fee, which is then added to the operator's own maintenance, sales, and channel costs, resulting in a naturally higher unit cost than if customers were to directly find model manufacturers.
Self-built mode (self-owned GPU + open-source weights): the Intelligent Computing Center purchases and deploys GPUs to run Qwen and the open-source version of DeepSeek,
The weights of models like ChatGLM are used for government and enterprise projects that require data localization and private deployment, which is the source of the "compliance premium". However, there are limitations: closed-source flagship models cannot obtain weights and can only run open-source and small vertical models; the ability to optimize engineering and vLLM inference is weak, and GPU utilization is not as high as that of major AI companies. Additionally, the costs of hardware depreciation, electricity, data centers, and maintenance personnel are stacked, making it uncertain whether running open-source models can be cost-competitive with leading manufacturers.
However, these three modes are still not the deepest root cause. What really drove the price up to a five-fold difference is the rigid cost accounting mechanism of state-owned enterprises: private cloud manufacturers can dynamically price, accurately calculate the models and traffic actually used by users, and even dare to "bet" that most users will use cheaper models most of the time. In contrast, as state-owned enterprises, telecom operators must price based on the worst-case scenario, assuming that users will fully utilize the most expensive large models, to ensure they do not incur losses - state-owned enterprises cannot lose money, as losses would trigger accountability. As a result, the three-tier packages are uniformly priced, with the same unit price: it is not based on your actual usage, but rather on the assumption that you might use the most expensive model to its full capacity.
This precisely illustrates that the essence of operators' token business is not technical capability, but rather a rigid price increase under the mechanism of compliance licenses and state-owned enterprise pricing - customers are buying a "non-outbound, compliant, and worst-case-scenario-priced model", rather than a "latest and strongest, usage-based-priced model". Operators have neither technical leadership, nor distribution rights, nor pricing flexibility, and their token business is a compliant niche market from start to finish.
Therefore, the true structure of operators' token business can be concluded: voice fee payment channels + hybrid supply mode (agency + self-built + joint venture) + compliant packaging + state-owned enterprise rigid pricing - low technical content, no pricing power, and no elasticity. The only moat is the operators' enterprise-level settlement capability - not only voice fee deduction, but also government and enterprise bill collection, unified settlement, and central enterprise procurement access, which is a set of capabilities that cloud vendors and model vendors lack. However, its ceiling is also limited: this settlement capability is essentially "channel capability" rather than "technical capability", and as model vendors deepen their cooperation with payment, financial, and government and enterprise service providers, the channel advantage will be gradually bypassed.
There are also objective resource constraints. Large models are bottomless pits for capital consumption, and for a central enterprise with annual revenues of tens of billions and single-digit profit margins, putting a large amount of funds into a general large model with an uncertain outcome is a commercial risk that is too high.
But keeping a distance from general large models does not mean keeping a distance from large models altogether. Building a domestic computing power base, implementing independent innovation and substitution, participating in the Eastern Data and Western Computing project, and constructing a national computing power network are strategic tasks assigned to operators by the nation, which also perfectly match their resource endowments. Central enterprise resources should be prioritized towards the construction of bases with high certainty.
A key distinction needs to be made here: not doing general large models does not mean not doing models at all. For scenarios such as communication operations, government affairs, urban governance, and industry customer service, operators can still lay out vertical small models and domain-specific models - these types of models are not comparable to C-end general large models, but rather serve their own local government and enterprise projects as a complementary capability to the algorithmic network foundation. Treating specialized models as complementary tools is different from going all out to create general large model products.
But another possibility should also be considered: once a vertical industry model becomes indispensable in a certain scenario, it may break away from being an "auxiliary" and grow into an independent high-value business - the communications network operation model, government hotline model, and anti-fraud control model all have independent paying intentions. However, this path requires data, scenarios, and talent, three things that operators currently lack. It is a long-term option, not a near-term reality. Currently positioning it as an auxiliary is a pragmatic starting point, but strategically, space should be left for it to grow independently, and the ceiling should not be nailed down from the beginning.
In the past, voice, text messaging, and data traffic were the core production factors of communication; with the advent of the AI era, computing power is expected to become the fourth factor. However, factors are not equal to products - carriers make money from voice, text messaging, and data traffic not because they produce the best instant messaging services, but because they control the measurement, scheduling, and distribution of these factors.
Following this logic, the profit focus of operators in the AI era should not be on C-end large model applications, but rather on the scheduling, distribution, and service of AI core elements. The positioning is not to develop general large models to compete with DeepSeek, Qianwen, and Doupai for the C-end market, but rather to focus on three things: global scheduling of computing power networks, domestic large model computing power base, and providing a "network + computing power" integrated base for large model manufacturers across the industry - acting as the "pipeline" for large models.
However, there is a major premise that must be faced: the logic that "elements ≠ products" holds true because operators have the distribution rights of elements. In the era of voice, SMS, and data traffic, operators were the gatekeepers - numbers, spectrum, base stations, and pipelines were all physical monopolies, and OTT services had to pay to pass through. The right to distribution is essentially a physical monopoly right.
With the advent of the AI era, this premise has been shaken. The right to produce tokens lies with model manufacturers, while the calling entrance (or distribution rights) is in the hands of cloud providers and application layers; even the underlying computing power has three competing routes, including cloud providers' self-built intelligent computing centers, local state-owned capital computing power platforms, and large model manufacturers' self-built facilities. Operators are now left with only the "road" (network) and a portion of the "parking lot" (computing power).
The more-than-fivefold price difference is ironclad evidence of this - what's expensive is the compliant license, not the technical capability. If one truly has distribution rights, they should be collecting distribution fees, rather than relying on high unit prices to force sales. High prices are a symptom of having lost distribution rights, not evidence of holding pricing power.
There are three key factors in commercializing a business, each with its own boundaries that must be clearly understood:
First, the B-end compliance premium. It provides CDN edge acceleration, dedicated lines, 5G slicing, and computing power clusters to government and enterprise customers. However, it's necessary to clarify that network slicing and QoS priority are subject to network neutrality constraints, and the space for large-scale sales of differentiated acceleration is limited. The true structural premium mainly comes from the risk premium brought by Zhongchuang adaptation, local data storage, Grade Protection, and confidential qualification - government and enterprise customers pay for security and compliance, and this type of premium is resistant to price wars. On the C-end side, only moderate opening of compliant scenarios such as game acceleration is allowed.
Second, localized delivery. The demand for domestic self-developed information technology from local government and enterprise customers has unique advantages in terms of operators' cloud and computing power, with many projects being procured through single-source procurement. The local government and enterprise teams accumulated over decades are valuable assets, and it would be extremely costly for cloud manufacturers to replicate a ground team that covers the entire country. However, it is necessary to acknowledge the shortcomings: local customer managers are skilled at relationship maintenance, but generally lack the ability to provide AI solution capabilities, and easily fall into the trap of "obtaining projects but failing to deliver". A feasible approach is for "operators to grasp leads and partner with ISV ecosystem partners to complete solution implementation".
Third, computing-network convergence. By integrating communication networks and computing power networks, a unified "dedicated line + edge computing cluster" solution is launched, multiplying the network with computing power to increase customer migration costs. Once a customer's IT architecture is built based on computing-network convergence, a strong binding is formed. This is a unique combined capability of operators that pure cloud vendors find difficult to replicate.
Learning from Mistakes: Avoiding the Commodification of Pipelined Products
The fate of pipelines is commoditization. The pain point of the first generation of OTT is not the disappearance of pipeline revenue - traffic revenue has been increasing - but rather that metered billing does not create stickiness: no differentiation → competing on price → price war → no increase in revenue despite increased traffic. Traffic has grown 100 times, but ARPU has remained stagnant.
Looking ahead to the AI era, if the "pipeline of large models" continues to charge based on token invocation volume and computing power duration, it is likely that in a few years, the surge in computing power demand and the continuous decline in unit revenue will repeat. The breakthrough does not lie in technological upgrades, but in activating channels. Operators have an underestimated offline force: hundreds of millions of stores, grid teams, government and enterprise customer managers, and SA partners. These channels should not just sell phone cards and renewal fees, but should be transformed into a distribution network for computing power and AI services.
This leads to three design principles:
First, the revenue-sharing ratio must be anchored to the concept of "channels being worth learning." The complexity of the mining machine serial number cards requires channels to invest in learning costs, and the revenue sharing must cover these costs, otherwise, stores will revert to selling serial number cards.
Second, it is divided into a long cycle. The one-time commission is changed to an annuity, forcing the channel to value renewals over one-time transactions.
Thirdly, it adopts a layered and segmented approach. Large government and enterprise projects are handled by customer managers and industry partners, while the sales of computing power to small, medium, and micro enterprises are entrusted to agents, with stores only responsible for lead generation and not in-depth sales of complex computing power.
There are still two organizational hurdles to overcome: the revenue-sharing mechanism spans the cloud division, network division, government and enterprise division, and channel division, involving performance evaluation and reallocation of interests, which faces significant internal resistance and requires advance clarification of performance evaluation and cross-departmental interests. At the product level, standardized packages (such as bundling "dedicated lines + edge computing power + security foundation") need to be established to reduce channel understanding and delivery costs, making them suitable for large-scale distribution.
Only in this way can the moat, which relies heavily on manpower and services, be truly difficult to replicate - internet giants pursue lightweight operations and will not maintain a ground team of millions, which is the most core barrier for operators.
Facing Risks: Local Markets Are Not Exclusive Territories
This line of thinking also faces external competition, and not just from major internet companies.
Alibaba Cloud and DingTalk have accelerated their localization efforts, deeply participating in the digitalization of government services in multiple regions, with many benchmark projects being implemented by internet companies - the local government and enterprise market has never been the exclusive domain of telecom operators. Telecom operators have two advantages: firstly, they have the backing of central enterprises, and in projects involving confidential information and critical infrastructure, "Alibaba is still a capitalized company" with limitations in terms of access; secondly, they have a deep and extensive channel network, while large factories can secure benchmark projects, they are unable to cover the long-tail market in cities and counties across the country.
However, it's crucial to stay vigilant: the backing of central enterprises is merely an entry ticket, not a permanent moat. If a company chooses to be complacent, failing to refine its products and enhance its AI solution capabilities, it will continue to lose benchmark projects, and relying solely on long-tail businesses will not be enough to support its operations in the AI era.
In addition to internet giants, there are two other types of competitors. Local state-owned computing power platforms are more direct competitors: they have the backing of local governments and have equivalent or even stronger competitiveness and operating capabilities in local government-enterprise projects - after all, they are the "favored sons" of local governments. AI large model manufacturers, on the other hand, build their own computing power centers and deploy edge nodes, replacing the computing power pipelines of operators.
It's also necessary to dispel an easy self-consoling illusion: compliance is just an entry ticket and does not equal high profit margins. Building GPU-based open-source model services for Zhonghua Xincai projects also requires bearing the costs of continuous optimization of open-source models, data annotation, talent recruitment, and version iteration - open-source does not mean free, and deployment does not mean delivery. Operators need to combine algorithm network integration, local services, and security compliance to form a comprehensive barrier, rather than simply competing on computing power prices.
The path for operators' OTT transformation can be summed up in one sentence: blocking the entrance is not feasible, and competing with general large models is not winnable; the correct approach is the "base + pipeline + channel revenue sharing" trilogy - solidifying the domestic algorithmic power base with national strategy to defend the basic market; defending high profit margins with compliance premiums and localized delivery; and using a layered channel revenue sharing mechanism to counter pipeline commodification.
The first time it was OTT, operators lost their value-added businesses; in the era of large models, if they only act as transparent pipelines collecting tolls, they will lose the entire AI era. Being a pipeline is not shameful, what's shameful is only knowing how to collect tolls. Operators need to become new-type pipelines with service capabilities, ecosystem capabilities, and local delivery capabilities - and also reclaim some of their distribution rights.
