Economic Observer Follow
2026-07-25 16:33

Economic Observer reporter Zheng Chenye
At the 2026 World Artificial Intelligence Conference (WAIC) held from July 17th to 20th, supernodes became the focus of the entire event. According to incomplete analysis by reporters, over 20 companies have released products or solutions related to supernodes.
As early as the Huawei Connect Conference in September 2025, the Ascend Atlas950 supernode was first released, but at that time, only product parameters and rendering images were seen by the outside world, and the real machine did not appear. In the past year, the actual performance of the 950 super node and whether the mass production progress has been pushed forward as scheduled have been the focus of inquiries from investors and industry professionals who are concerned about domestic computing power.
It doesn't matter if it's 100% expensive, it doesn't matter if it's 200% expensive
In July of this year, a recorded transcript of DeepSeek founder Liang Wenfeng's speech at an investor exchange held two months ago began to circulate in the investment community. In this nearly four hour communication record, Liang Wenfeng made the above statement when talking about the price difference between Huawei hypernodes and Nvidia products.
Liang Wenfeng also expressed his views on DeepSeek's open-source strategy, AGI (General Artificial Intelligence) technology roadmap, API (Application Programming Interface, a technical channel for external users to call AI models) pricing logic, as well as the gap in computing power between China and the United States and the prospects for domestic chip substitution. The evaluation of domestic computing power has sparked widespread discussion on the internet.
A source close to DeepSeek investors confirmed the authenticity of the manuscript to Economic Observer reporters. On July 24th, the reporter also verified the authenticity of this audio transcript to DeepSeek through phone and email, but as of the time of publication, there has been no response.
According to Liang Wenfeng, the Huawei 950 supernode can replace similar supernode systems such as Nvidia GB200NVL72 at the task level. The former can do all the tasks that the latter can do, with consistent latency performance. The cost is that the computing power of four Huawei cards is about equivalent to one Nvidia card, which lags behind the chip iteration pace by about two years.
Hypernodes have been the most popular product direction in the field of domestic computing power in recent years. It connects dozens, even thousands, and tens of thousands of AI chips into a tightly coordinated computing system through a dedicated high-speed bus. The communication bandwidth between chips can reach hundreds of GB or even TB level per second, and the latency is as low as nanoseconds, making the entire cabinet or even multiple cabinet machines behave like the same supercomputer at the software level. Compared to traditional server clusters that use Ethernet cables to connect multiple machines, the communication speed between chips within a supernode is close to the level of chips, allowing multiple cards to work together like one card.
As the parameters of large models move from billions to trillions or even trillions, traditional 8-card servers (one machine with 8 AI chips) can no longer meet the communication needs of training and inference, and supernodes have become a must-have for various computing power manufacturers.
From Huawei's commercial use of over 750 sets of Ascend 384 supernodes, to the planned launch of Atlas950 in the fourth quarter of this year (which can support up to 8192 cards with full configuration), to the recently completed 100000 card Shuguang 8000 Supercluster by Zhongke Shuguang, domestic manufacturers have already answered the question of whether there is a domestic alternative solution for top computing products.
Liang Wenfeng, as the founder of a leading domestic model company, publicly gave the evaluation of "can be replaced", which to some extent dispelled the market's doubts about the actual availability of domestic super nodes. But the question of how to implement the business model of domestic super nodes has just surfaced.
Previously, almost all super nodes were delivered to customers in the form of whole cabinet hardware. A set of systems can cost tens of millions of yuan at least, and hundreds of millions of yuan at most. Customers need to build their own computer rooms and operate and maintain their own. Only large Internet enterprises, telecom operators and national intelligent computing centers can afford to deploy them.
But this situation is changing. On July 23rd, Alibaba Cloud announced that the Lingjun Zhenwu M890 super node cloud computing power unit has successfully adapted to the 2.4 trillion parameter Qwen3.8 model and launched inference services on the Bailian platform.
As a result, M890 has become the first super node product in China to stably run over two trillion parameter large models on a public cloud (an open, pay as you go cloud computing platform for all users, different from private clouds built by enterprises). Users do not need to purchase hardware themselves. They can apply for a 64 card fully interconnected cloud computing power unit on the Alibaba Cloud console, which is paid by the hour at a rate of 120 yuan per hour and released upon use.
In the same period, the 100000 card cluster operated by Zhongke Shuguang (603019. SH) connected to the national supercomputing Internet and charged users nationwide according to the usage; Huawei Cloud announced a comprehensive shift from selling cloud servers to billing based on tokens.
During the interview, the reporter learned that the technical feasibility of domestic hypernodes has been basically verified in the past two years, and the differentiation of business models has just begun. The delivery method of hypernodes is evolving from a single full container hardware sales to multiple paths - selling full containers, selling computing power, and selling keywords. Customers of different sizes are starting to have different choices.
From 'Made' to 'Used'
During the WAIC period, the highly anticipated product, the Ascend Atlas950 Super Node, was finally exhibited for the first time in its actual form, becoming one of the most highly anticipated exhibits in the exhibition.
The reporter learned from the relevant person in charge of Huawei that the 1024 card version exhibited on site is based on a single cabinet of 64 cards as the basic unit, which can provide 1EFLOPS (billions of floating point operations per second) FP8 (a commonly used low precision data format in AI computing) computing power, equipped with 256TB of global unified memory, and the round-trip communication delay between chips is as low as 3 microseconds. Huawei has also developed an interconnection protocol called Lingqu 2.0 for this system, which reduces the cross card communication latency from the traditional network's 2 microseconds to 200 nanoseconds, equivalent to allowing thousands of cards to communicate in real-time using the same "language". In addition, the fully equipped 8192 card version is planned to be launched in bulk in the fourth quarter of this year, and Huawei officials have stated that multiple top customers have already completed their orders.
According to the introduction of the relevant person in charge of Huawei, Shengteng 384 super node is the only super node in China that has trained the SOTA (currently the best) model. At present, it has more than 750 sets of super nodes for commercial use worldwide, covering the Internet, operators, finance, education, medical and other industries.
Huawei also showcased the Atlas850E air-cooled super node at the same time, which adopts self-developed VCE (steam chamber equalization) phase change heat dissipation technology and can be directly deployed in standard air-cooled machine rooms without the need for liquid cooling transformation. It is mainly aimed at inference scenarios, reducing the deployment threshold for enterprises without liquid cooling conditions.
What Zhongke Shuguang wants to verify is another dimension of the problem - when the scale of domestic chips expands from kilocards to 100000 cards, can the system still run stably for a long time?
On July 10, Dawning 8000 (Peak) 100000 card super cluster was completed and put into operation in Zhengzhou, and synchronously connected to the national super computing Internet, becoming the first national 100000 card AI super cluster in China. This system is equipped with the DCU chip of Haiguang Information (688041. SH), combined with Shuguang's self-developed high-speed interconnect network.
Bu Jingde, Chief Engineer of High end Computing at Zhongke Shuguang, told Economic Observer that increasing the number of machines from 10000 to 100000 is not a simple matter. The power consumption impact on the power grid when the 100000 card system is instantly started far exceeds that of the 10000 card system. The entire system has 3 billion components, and any small fluctuation can affect global stability. Network fault recovery must be done in seconds.
Heat dissipation is also an engineering challenge. Cao Zhennan, Vice President of Zhongke Shuguang, introduced that Shuguang 8000 adopts a three-layer circulating cooling scheme. The innermost layer uses special liquid evaporation to remove chip heat, and the outermost layer is directly connected to local lake water in Zhengzhou for natural heat dissipation. The entire host system does not require traditional air conditioning, and PUE (a measure of data center energy utilization efficiency, the closer the value is to 1, the less energy waste) is controlled at an extremely low level.
According to the introduction from Zhongke Shuguang, in the first week of its launch, Shuguang 8000 was running at full load, processing over 150000 jobs per day, with a peak of over 500000. If all of this computing power is used for large-scale model inference, it can simultaneously support over a million people asking AI questions and receiving real-time responses.
In an interview with reporters, Tan Guangming, Secretary General of the High Performance Computing Committee of the Chinese Computer Society (CCF), also mentioned that the value of the Shuguang 8000 for AI for Sci ence (using AI for scientific research) is particularly prominent. It provides a large-scale real verification platform for domestic research on super intelligent fusion that was not previously available.
Beyond Huawei and Shuguang, the popularity of supernodes is spreading throughout the entire industry chain. According to incomplete statistics from reporters, over 20 companies have released hypernode related solutions on this WAIC.
For example, ZTE Corporation (000063. SZ), in collaboration with multiple domestic GPU manufacturers such as Weiren Technology, Muxi Corporation (688802.SH), Suiyuan Technology, and Tiantian Zhixin, has launched Matrix supernodes that support multi chip hybrid insertion; New H3C (specializing in enterprise level IT infrastructure) under Tsinghua Unigroup (000938. SZ) has released the UniPoDS80000 series, which can be flexibly expanded from 32 cards to 16384 cards; Muxi Corporation has released the Xi Jing S600 with 64 cards as the basic specification.
Based on the above situation, it can be found that there are three technological routes emerging around supernodes: the full stack self-developed route represented by Huawei, which is based entirely on Huawei's self-developed chips, interconnection protocols, and software stacks; The compatibility route represented by ZTE and New H3C supports GPU hybrid deployment from multiple manufacturers; The joint innovation route taken by chip manufacturers such as Muxi and Suiyuan in collaboration with whole machine manufacturers.
Not only is hardware advancing, but the validation of the software ecosystem is also being carried out simultaneously.
Zhang Jianzhong, founder of Moore Thread (688795. SH), proposed the concept framework of the "era of lexical elements" at the WAIC sub forum, and divided computing infrastructure into three categories: model training factories, lexical element production factories, and intelligent agent production factories.
According to Zhang Jianzhong's introduction, Moore Thread has completed the complete training of a 236 billion parameter MoE (Hybrid Expert, an architecture that allows large models to activate only some parameters to improve efficiency) model from scratch on a 10000 card cluster built on the self-developed GPU chip MTTS5000. The training quality can benchmark the international mainstream level.
This is also the first time that a domestic GPU manufacturer has publicly demonstrated the results of such a large-scale model training practice, which to some extent responds to the market's doubts about whether domestic chips can train large models.
During this WAIC, Pengcheng Laboratory and the Global Computing Alliance jointly released the "White Paper on the Definition and Practice of Hypernodes", which for the first time provided a formal definition of hypernodes. They are physically composed of multiple computing nodes tightly connected through efficient interconnection protocols, with unified memory addressing capabilities across physical nodes, and logically possessing the characteristics of "one computer".
The white paper also clarifies the three core technical features of supernodes, namely memory semantic interconnection, unified memory addressing, and ultra-low latency and ultra large bandwidth. Among them, "unified memory addressing" is regarded as a key threshold for determining whether a system belongs to a true supernode - only when all chips can read and write data as if sharing the same block of memory, can communication bottlenecks between chips be eliminated. Otherwise, no matter how large the system scale is, computational efficiency cannot be improved. The white paper proposes that in the next two to three years, the main focus of global AI infrastructure competition will shift from single chip performance competition to system level engineering capability competition.
Liang Wenfeng's view on the domestic super node ecosystem is more positive than industry consensus. At the aforementioned investor exchange meeting, he stated that DeepSeek trained the V3 model almost without relying on Nvidia's software ecosystem, and instead completed all the work based on its self-developed TileLang advanced compiler.
In his view, the moat of the NVIDIA CUDA software ecosystem is rapidly crumbling for three reasons: AI itself can write code to rebuild the ecosystem, high-level compiled languages such as TileLang have significantly lowered the development threshold, and GPU architecture coupling derived from game cards is being unbound.
The hardware and ecosystem of domestically produced AI chips no longer have fundamental obstacles, the only bottleneck is production capacity. ”Liang Wenfeng also disclosed at the above investor exchange meeting that Huawei's current capacity for DeepSeek is about 16000 cards, while that of Internet giants is more than 100000.
Liang Wenfeng also summarized the gap between China and the United States in the field of AI as "lagging behind by one to two years, using only one twentieth of its computing power". In his view, this gap is mainly caused by unequal computing power resources, rather than differences in talent or technological routes.
Three paths to sell computing power
Hypernodes are not just simple stacks of AI chips, a complete system also includes high-speed switching networks, liquid cooled cooling, high-voltage power supply, and system engineering to integrate these components into the entire cabinet.
Taking Nvidia as a reference, its previous generation GB200NVL72 (Nvidia's 72 card full container super node launched in 2024) was priced at about $3 million (approximately RMB 21 million), and the material cost of the new generation VeraRubin NVL72 produced this year has further increased to $7.8 million per container.
There is no publicly available standard selling price for domestic supernodes, but China Mobile's centralized procurement this year provides a reference -776 sets of supernode computing node equipment, with quotes from five manufacturers all exceeding 2 billion yuan. However, China Mobile purchases 384 card computing nodes, not full container systems. In actual deployment, it also requires supporting exchange cabinets, optical modules, and liquid cooling infrastructure. The total cost of landing a complete 384 card super node is much higher than this.
For telecom operators and Internet giants, this is an affordable investment in infrastructure, but AI Computing Power has far more customers than these big buyers. University laboratories, AI startups, and small and medium-sized development teams also require top-notch computing power, but their usage is limited and they cannot afford the high one-time investment.
The gap in demand scale is driving the differentiation of delivery methods for hypernodes into at least three paths, targeting customers of different sizes and corresponding to different profit logics.
The first path is to sell full container hardware.
Huawei Ascend supernodes currently mainly adopt this model, delivering them to customers in full container form, and customers deploy and operate them in their own data centers. The buyers are mainly three major telecom operators, large Internet enterprises and national intelligent computing centers. These institutions have data centers, operation and maintenance teams, and continuous computing power needs. Buying hardware and building their own is the most economical choice.
For example, China Mobile's centralized procurement of supernodes launched in March this year is currently the largest single supernode order in public information, and also the first large-scale procurement of supernode equipment by the three major telecommunications operators at the group level. This centralized procurement involves a total of 776 sets of supernode computing node devices and 6208 AI acceleration cards, fully targeting the Huawei CANN (Ascend Computing Architecture) ecosystem. According to the announcement of the winning bidders, five manufacturers including Henan Kunlun, Changjiang Computing, Huakun Zhenyu, Baode Computer, and Huaqi Intelligence have been shortlisted, with bids ranging from 2.059 billion to 2.07 billion yuan, with a difference of only 11 million yuan between the highest and lowest bids.
A long-term investor who has been paying attention to the domestic computing power field told reporters that this extremely close quotation reflects that the profit of the hardware link has been squeezed very thin, but the initial hardware delivery is not the main source of profit for this model. The hypernode integrates modules such as liquid cooling, high-speed exchange, and power management into the delivery scope. The gross profit margin for subsequent system expansion, maintenance, and on-site services is usually over 40%, far higher than the initial hardware.
According to public information, the three major telecommunications operators will have a total computing capital expenditure of over 80 billion yuan in 2026, of which China Mobile's computing network investment will be 37.8 billion yuan, a year-on-year increase of 62.4%.
It is worth mentioning that Liang Wenfeng also revealed in the aforementioned communication that DeepSeek's computing power cluster is all self built, and the investment return of buying cards is much higher than keeping funds in the account. I can turn money into cards and recover the cost in ten months. I will definitely buy as much as I can, "said Liang Wenfeng.
The second path is self built and self operated.
The Shuguang 8000 is following this path. The 100000 card cluster is not sold to a specific customer, but is operated by Zhongke Shuguang itself. This mode relies on the national supercomputing Internet. The supercomputing Internet is a nationwide computing power scheduling network promoted by the Ministry of Science and Technology. The goal is to connect more than 30 supercomputing centers and intelligent computing centers across the country, so that users can use computing power on demand by placing orders on the platform, regardless of which physical machine room the computing power comes from.
On July 13 of this year, supercomputing Internet launched its core node in Zhengzhou, positioning itself as the central dispatching hub of national computing power. The registered users of the platform have exceeded 1.4 million, and the dual track pricing of "inclusive+market" is adopted. Scientific research users can obtain free computing power quota by registering, while commercial customers are charged according to market price.
The reporter learned from Zhongke Shuguang that Shuguang 8000 was directly connected to this network after its completion in Zhengzhou, providing computing power services to scientific research institutions and enterprises across the country. The pricing ranges from 1.2 yuan to 3.5 yuan per card hour (the billing unit for one hour of AI chip operation), and scientific research institutions can enjoy up to 50% universal subsidies, with a minimum subsidy of less than 1 yuan per card hour after the subsidy. In the first week of launch, the demand for computing power waiting in line has exceeded three times the total size of the cluster.
The profit logic of this path is completely different from selling hardware. The aforementioned investor pointed out that selling hardware is a project-based income, which ends when one server is sold. The gross profit margin is around 20% and is affected by fluctuations in customer capital expenditure cycles; Operating computing power is a sustainable income, and as long as the cluster utilization rate remains at a high level, the comprehensive gross profit margin can reach over 40%.
According to the 2025 annual report of Zhongke Shuguang, the proportion of computing power service revenue has increased from 12% in 2023 to 37%, becoming the fastest-growing business sector.
Liu Wenshu, Chief Analyst of Computer at Zheshang Securities, said that if computing power manufacturers can shift their profit model from one-time hardware sales to continuous computing power services, they can obtain more sustainable cash flow. At the same time, value-added services such as cluster operation and maintenance hosting, and improving the efficiency of keyword usage are also expected to further increase gross profit margins.
The third path is the public cloud.
The Alibaba Cloud M890 supernode cloud computing power unit has turned supernodes into on-demand cloud services. The 64 card fully interconnected cloud computing power unit costs 120 yuan per hour, with an average of less than 2 yuan per card hour. It also supports billing based on task packages and keyword call volume.
During his time at WAIC, Wu Jiesheng, the head of Alibaba Cloud's elastic computing product line, stated that when a trillion parameter large model requires more efficient inference, and training tasks often span hundreds or thousands of cards, the real answer is to integrate chips, servers, networks, and storage into a whole, providing an "AI supercomputer" that can train, reason, and be used out of the box.
In the design of Alibaba Cloud, users do not need to buy hardware or build a data center themselves. They only need to open it on the cloud platform to directly obtain a super node computing environment to run large models.
Huawei Cloud is also evolving in a similar direction. Huawei Cloud CEO Zhang Ping'an previously publicly stated that Huawei Cloud has fully shifted from the IaaS (Infrastructure as a Service) model of selling cloud servers to the MaaS (Model as a Service) model of selling keywords. Public data shows that Huawei Cloud's CloudMatrix384 (the product name of Huawei Cloud's super node cloud service), based on the Ascend 384 super node, has a cost of about 1.8 yuan per million words, which has strong price competitiveness among domestic computing power solutions.
Of course, there is also cross infiltration between the above-mentioned paths. For example, Huawei serves both telecom operators through full container sales and small and medium-sized customers through Huawei Cloud's MaaS model; Shuguang not only operates its own computing power cluster, but also sells hypernode hardware and overall solutions to other intelligent computing centers.
In other words, different paths have formed a hierarchical structure for different customers: large operators and Internet enterprises have computer rooms, teams, and continuous demand, and it is most economical to buy a whole counter and build it by themselves; Medium sized enterprises and research institutions have significant fluctuations in computing power demand, and renting on demand can avoid heavy asset investment and idle equipment; Small and medium-sized developers can pay by the hour on public clouds, and for a few hundred yuan, they can use the top computing power that was previously only accessible to large companies, which is flexible and convenient.
Evolution towards a universal computing power base that integrates training and promotion
Whether computing power can continue to operate efficiently is becoming the core variable that determines whether this market can deliver commercial returns.
The first week full load of the Dawn 8000 is good news, but not all intelligent computing centers have such demand support. According to the data released by the Ministry of Industry and Information Technology at the press conference of the State Council Information Office on July 20th, as of the end of June this year, the scale of intelligent computing power in China has reached 2185EEFOPS, with the eastern region accounting for 55.9%, the western region accounting for 32.6%, and the central region accounting for only 10.6%. Regarding this, an investor who has long been concerned about the computing power field told reporters that the eastern intelligent computing centers are generally operating at high loads, while some intelligent computing centers in the central and western regions are facing a shortage of customers and low utilization rates after completion. The higher the computing power and density, the greater the cost pressure of idle.
The operating costs of electricity, operation and maintenance, bandwidth, depreciation, and other expenses in the intelligent computing center are rigid expenditures. The larger the cluster size, the higher the fixed daily expenses. If the utilization rate is below a certain threshold, it means continuous losses. Especially for enterprises that have just transformed into computing power operators, how to continuously find enough customers and maintain high utilization rates after completion is a more difficult question to answer than technology.
The national supercomputing Internet is trying to alleviate this geographical mismatch between supply and demand.
The Zhengzhou core node, which was launched on July 13th, is positioned as the central dispatch hub for computing power in China. Users can access the platform to access computing resources across the country as needed, regardless of which city they are in. During the interview, the reporter learned that it is expected that by the end of 2026, the annual transaction volume of national supercomputing Internet will exceed 10 billion yuan. In addition, Henan Province, which has established a central dispatch hub, has also established a computing voucher mechanism at the provincial level, with an annual distribution of no more than 50 million yuan, to reduce the computing costs of small and medium-sized enterprises.
From the demand side, new application scenarios are also creating incremental demand.
Du Xiawei, Assistant to the President of Haiguang Information and General Manager of the Intelligent Computing Product Department, mentioned in an interview with reporters that the rise of AIAgents (AI systems that can autonomously plan and execute complex tasks) is changing the structure of computing power demand. In the past, AI mainly used GPUs for model inference, with CPUs only playing an auxiliary scheduling role. However, when executing a complex task, agents require CPUs to complete a large amount of work such as task orchestration, tool calling, and cache management. The ratio of CPU and GPU computing power is changing from 1:8 in the past to 1:1 or even higher. This means that the demand for computing infrastructure is not decreasing, but becoming more diverse.
Li Ran, Chief Engineer of Zhongke Shuguang AI Industry, also believes that the load in the Agent era is long links and high consumption, requiring the overall coordination of computing, storage, network, and software. Shuguang's response is to create a full link inference acceleration solution around supernodes, optimizing end-to-end from the underlying operator library to the deployment of upper level models.
In his view, supernodes are evolving from "training specific facilities" to "universal computing power bases that integrate training and promotion", and their application scenarios have expanded from large model training to inference Agent、 More fields such as scientific computing.
Pan Helin, a member of the Information and Communication Economy Expert Committee of the Ministry of Industry and Information Technology, told Economic Observer reporters that the essential difference between supernodes and traditional GPU computing power clusters lies in whether they have achieved unified memory addressing and vertical expansion level interconnection. Traditional GPU computing power clusters adopt horizontal scaling, with high latency and low bandwidth, while supernodes can build unified memory addressing, with latency as low as microseconds and bandwidth up to TB per second, and can directly access cross node resources.
Simply put, supernodes make up for the shortcomings of domestic computing power single-chip performance, making it possible for us to catch up with the world's advanced computing power level, "said Pan Helin.
However, in Pan Helin's view, in order for domestic super nodes to truly achieve large-scale commercial use, they still need to overcome two obstacles. One is ecological adaptation. Currently, there are different connection protocols among the hypernodes of various manufacturers. How to unify these protocols to fully utilize the efficiency of interconnection, and at the same time, adapt mainstream models such as DeepSeekV4, Qwen, GLM, etc. to reduce the adaptation threshold for developers and avoid duplicate development, is still a question that the industry needs to answer; The second is to optimize costs while improving system reliability. The simultaneous access of thousands or even tens of thousands of cards to supernodes may lead to a decrease in overall system stability, requiring the improvement of supporting mechanisms such as fault response and self-healing. The reduction in computing power costs depends on the scale of nodes and the continuous drop in single node costs, in order to attract more AI large model users to access supernode networks.
Yang Lin, Chief Analyst of Guotai Haitong Securities Computer, told reporters that the core paradigm of computing power competition has shifted from "single card peak" to "system efficiency", and value distribution is fully shifting from "selling cards" to "selling systems and capabilities". Optical interconnection, liquid cooling, network solutions, system integration and scheduling software and other links will receive higher value weight.
According to Huatai Securities' calculations in relevant research reports, the market size of China's super node architecture is expected to reach 341.4 billion yuan in 2028, with a compound annual growth rate of 194% from 2026 to 2028.

The front air spring has hidden dangers, and Xiaopeng X9 has initiated a recall

Test drive Huawei Qiankun ADS 5.0: WEWA 2.0 architecture first test, will be gradually installed on Huawei's models

The car owner's self driving new energy vehicle's overseas car infotainment system has been remotely locked for over 30 hours, and the brand involved has responded by triggering the built-in security protection mechanism of the car infotainment system, which is a commonly used anti-theft and damage prevention measure in the industry