
In 2026, the AI industry will begin to rethink a fundamental question: what should the foundation of computing power look like.
The parameters of large models are rapidly moving towards 10 trillion, and intelligent agents are moving from hourly level work to monthly level. China's daily inference tokens (word elements, the basic unit for processing information in large models) have reached 500 trillion, and will reach millions of trillion by 2030; End side intelligence is also rapidly enhancing, with mobile phone models evolving from 3B in 2024 to 30B today, and moving towards the 100B-level in the future. The demand is growing by orders of magnitude, and improving chip performance alone is no longer enough. The supply of computing power needs to be redesigned from the bottom.
The reason why this issue is urgent is that the division of labor in the AI industry is moving towards deeper waters. For those who create models, applications, and infrastructure, their respective boundaries need to be clear. Especially for the side that makes the base, if they are involved in both the model and the application, it will be difficult for the upper level enterprises to build their business on it with confidence.
On September 17th, the Huawei Connect Conference 2026 opened in Shanghai. Huawei Vice Chairman and Rotating Chairman Wang Tao established silicon-based black soil as the core expression of Huawei's AI strategy in his keynote speech, focusing on AI infrastructure, serving as the foundation of computing power, and leaving the prosperity of applications and data to partners.
On site, Huawei also released the industry's first supernode Ascend 960 using NPO (Near Package Optics) technology, and announced the roadmap for Ascend chips in the following years.
On that day, Wang Tao had in-depth exchanges with media such as the Economic Observer, and further elaborated on topics such as the technological boundaries between supernodes and clusters, the industrial prospects of near packaged optics, and the open path of computing power ecology.
The basic unit of calculation has changed
For decades, the basic unit of data centers has been servers, with procurement calculated on a per unit basis, deployment pushed forward on a per unit basis, expansion stacked on a per unit basis, and the way to expand computing power scale is to buy more servers and connect them to the same network.
The continuous iteration of AI big models has hit the ceiling of this logic that has been used for many years.
Currently, all top ranked SOTA (State of the Art) models have started with parameter scales at least in the trillion level, and will move towards the 10 trillion level by 2027.
Super node clusters with over 100000 cards are becoming the basic configuration of AI infrastructure. However, under traditional server architecture, communication within the cluster consumes over 40% of the training time, and the bandwidth between servers is much lower than within the servers. Although tens of thousands of cards are in the cluster, a considerable amount of time is not spent on computation and waiting for data.
An industry chain insider told reporters that in a current super large cluster with a scale of 100000 cards, the industry MFU (Model Floating Point Computing Utilization, measuring the degree to which computing power is actually used for computation) often accounts for less than 30%, which means that 70% of the time is spent in a state of failure and fault repair.
Buying tens of thousands of cards, 70% of the computing power is idle, which is difficult to calculate for any enterprise that invests in cutting-edge large model training.
Wang Tao mentioned that Huawei is taking a different innovation path, utilizing a brand new system architecture and adopting a "super node+cluster" approach to form a high-speed collaborative computing power system with multiple super nodes, in order to meet the training and inference needs of large-scale models.
The so-called supernode refers to a computing system that is physically composed of multiple computing nodes tightly connected through efficient interconnection protocols, has the ability to unify memory addressing across physical nodes, and logically presents itself as a single computer.
This year, Pengcheng Laboratory and the Global Computing Consortium (GCC) jointly defined supernodes with 38 units and included them in industry standards. Gao Wen, an academician of the CAE Member, previously publicly pointed out that super nodes have become a new paradigm for AI infrastructure construction.
Thousands of cards, logically speaking, are a computer.
Wang Tao pointed out that the memory across physical nodes within a supernode can be addressed uniformly, and any AI chip can access the entire memory sharing pool within the supernode, which is called a supernode.
Previously, the industry used a broad term for supernodes, but after the definition was written into industry standards, supernodes now have a unified measurement standard.
The industrial significance of unified memory addressing lies in the fact that communication through network protocols must be carried out packet by packet under traditional architectures, and no longer needs to occur within supernodes. The saved communication time can be directly converted into computing power output. According to relevant simulation data, the MFU of a 100000 card cluster composed of 4K supernodes is 2.75 times higher than that of an 8-card server cluster of the same scale. For every 1 percentage point increase in MFU, the training period of the 100000 card cluster's large model can be shortened by about three days.
In monthly training competitions, the difference in system efficiency is the difference in model iteration speed.
The benchmark for competition has shifted from single card peak computing power to system efficiency. The entry threshold has also changed accordingly, and any shortcomings in power supply, heat dissipation, interconnection, and rapid recovery of faults will lower the overall availability.
Wang Tao told reporters that it is not enough to make a good chip, nor is it enough to make a good server. Such a large-scale computing system ultimately requires complex multidisciplinary system engineering capabilities to deliver a highly available computing system to customers, which is valuable.
Huawei has long-term and large-scale research and development investment in every field such as communication technology, optical modules, computing power, and systems, and has the best experts in each field.
The difficulty of a supernode lies in its simultaneous requirements for chips, interconnects, heat dissipation, power supply, and software. Any weakness in any field can become a bottleneck for the entire system. Leading in a single field cannot solve the problem, and the accumulation of multiple fields at the same time reaches a certain level. Only when the system level innovation meets the conditions.
In addition, when the bottleneck of computing power shifts from computation itself to the connection between chips, communication capability becomes the core variable of computing power competition.
Interconnection protocols, networking architecture, and optoelectronic conversion, which used to be topics of communication engineering, now directly determine the efficiency limit of a computing power system. More than 30 years of communication accumulation have new applications in this era.
The scale of supernodes has evolved from 384 and 1024 cards to 4096 cards, with each generation enhancing the unified memory addressing capability of corresponding chips.
With the continuous elongation of intelligent agent tasks and the rapid expansion of model context, the form of computing power required for training and inference has expanded from a single intelligent computing supernode to a complex system. Intelligent computing supernodes are responsible for training and inference, while Kunpeng Tong computing supernodes are designed for general computing and intelligent agent operation scenarios. AI memory storage provides PB level caching capabilities for long task inference. Each subsystem is connected via Lingqu Unified Bus (Huawei's self-developed supernode unified interconnection protocol), and multiple interconnection protocols are unified into one, greatly reducing the cost of protocol conversion.
Furthermore, the supernode cluster can support a maximum of one million cards.
The expansion space in technology has been opened up, and commercial deployment is also following suit.
Up to now, more than 1000 sets of Shengteng 910C super nodes have been deployed, and Shengteng 950 super nodes have also been commercially used on a large scale. Its customers cover more than 370 enterprises such as the Internet, operators, finance, government affairs, and manufacturing. It is also the only super node in China that has trained the SOTA model.
The supply of computing power is no longer solely dependent on the performance improvement of a single chip. In the era of large model parameters moving towards 10 trillion, this path centered on system architecture innovation provides sustainable computing power support for China's AI industry and another scale validated option for global AI infrastructure.
Light has entered the computing system
Inside the supernode, the connections between chips are approaching the limit of copper.
After the inter chip interconnection speed entered the 224G era, the effective transmission distance of electrical signals was shortened to less than two meters, and the budget for single board loss was only about twenty dB (decibels, a unit measuring the degree of signal attenuation). And for a super node with a scale of thousands of cards, high-speed copper cables often count in the tens of thousands, and the burden of wiring, heat dissipation, and reliability increases synchronously with the scale.
The growth of system scale is endless, but the physical limit of copper will not move, making the replacement of interconnection methods inevitable.
The consensus of the entire industry is to bring light as close as possible to the computing chip, complete electro-optical conversion near the chip, and use light for long-distance and low loss connections that copper cannot achieve, such as in cabinet electricity and out of cabinet light.
Wang Tao pointed out that as the scale of supernodes continues to grow, it is inevitable to move from traditional pluggable optical modules to NPOs, which means installing the optical engine at a distance of only a few centimeters from the computing chip to complete the electro-optical conversion nearby.
The product that first paved this path appeared in China.
Huawei has introduced over 30 years of accumulated optical communication technology into supernodes and launched the NPO optical engine Hi ONE.

Wang Tao mentioned that Hi ONE has three industry firsts and is the first large-scale commercial NPO product in the industry. Currently, there are no similar products in the industry that have entered large-scale commercial use; Single engine 7.2T (transmitting 7.2 trillion bits per second), equivalent to integrating 36 200G channels into one module, is currently the industry's largest transmission capacity; It is also the only NPO product that implements a built-in light source.
Integrating the light source into the engine not only tests the optical design, but also the synergy between materials, processes, and manufacturing. The manufacturing yield alone has undergone several years of iterative efforts, which is one of the reasons why similar solutions in the industry have been stuck in the conceptual stage for a long time.
The Ascend 960 supernode replaces the originally required 48000 800G optical modules with 5500 Hi ONE, reducing power consumption by over 550 kW, doubling the system's fault free running time, and achieving an availability of 99.8%. The number of interconnected components is compressed by nearly an order of magnitude, and the base of failure probability is synchronously compressed, thus rewriting the reliability logic of ultra large scale systems.
This technological path follows the innovative direction of the Tao's law, which compresses the time and path of signal transmission beyond the miniaturization of geometric dimensions, and can also bring intergenerational improvements in system performance.
The CPO (co packaged optics) solution, which is also being explored in the industry, still faces industrialization challenges in terms of packaging yield and reliability. Near packaged optics is a more mature and feasible solution at this stage.
In the past, optical communication solved the transmission of information across cities and continents; Today, the same technological accumulation is used for interconnection between chips ranging from several meters to tens of meters. When light enters computing systems, the two technological lineages of optical communication and computing, which have evolved for decades, have merged on AI infrastructure.
After decades of accumulation, China's optical communication industry has entered the main battlefield of computing power competition for the first time.
Huawei has put forward a proposal for the establishment of an NPO standard in the global optical interconnect standard organization OIF and received industry response. The OPEN NPO multi-source protocol, which was initiated in collaboration with more than 20 industry partners, is also being promoted.
Standardization means that more manufacturers can produce compatible NPO products, and the overall cost of the industry decreases as the scale expands. A technology that is first implemented becomes a shared capability for the entire industry.
With the gradual establishment of the standard system, this technology direction, which was first completed by Chinese enterprises for large-scale commercial use, is increasingly becoming a common option for the global next-generation AI infrastructure.
The origin of internet standards often indicates where the next round of industrial innovation will revolve around.
The intelligent world requires silicon-based black soil
Silicon based black soil is not a new concept.
In 2018, Black Earth was proposed as a platform strategy, and Huawei positioned itself as the foundational platform for the intelligent world, leaving the prosperity of applications and data to its partners. After years of evolution, its connotation has expanded to an end-to-end computing power base covering cloud, edge, and end.

Wang Tao told reporters that Huawei's silicon-based black soil is an end-to-end AI computing power base that spans edge, end, and cloud sides. There are both large-scale computing power clusters used for training and efficient reasoning, and in small IoT terminals at the end, it should also fully move towards intelligence.
From training trillion parameter models in large-scale clusters, to smartphones in pockets, cars on the road, and sensors in factories, Huawei provides complete solutions from micro computing power to super computing power for different scenarios. The intelligent world is not just about the intelligence of data centers.
The positioning of the base delineates the boundaries of the business. The core of Huawei's AI strategy is computing power, insisting on hardware monetization and converging business at the bottom of industry stratification.
Wang Tao mentioned that the essence of insisting on hardware monetization is to constrain Huawei's own business boundaries, with 70% of R&D investment in software but not relying on software monetization. Every enterprise grinds its own tofu, and Huawei's tofu is the foundation of AI computing power, serving as the foundation and backing for large model enterprises. Guest appearances do not have long-term competitiveness.
In his opinion, the new generation of AI companies in China are very excellent, and Huawei's positioning is to support them well. Huawei is willing to provide two flexible computing models: hardware and cloud. Whether it is an Internet leader or an innovative newcomer, it also includes small and medium-sized enterprises. The core of Huawei is to provide computing rather than applications.
For large model enterprises, training and deploying models on a foundation that does not compete with them for business increases the certainty of long-term investment.
After the boundaries are clear, the Shengteng ecosystem begins to accelerate its growth.
Wang Tao mentioned that after eight years of effort, Shengteng Ecology has basically crossed the turning point.
CANN (Ascend's basic software platform, through which developers call Ascend's computing power) code is fully open sourced. On the platform of the China Open Atomic Open Source Foundation, CANN has become the first active community, with external developers accounting for 61%, surpassing internal developers for the first time. The monthly active developers in the community have exceeded 5200, and there are over 40 models trained based on Ascend and CANN native.
There have been cases in the community where customers have self optimized operators based on open source code that outperform the original version, which is difficult to occur in ecosystems that were previously maintained unilaterally by vendors. Open source allows R&D teams to face the scrutiny of global developers, resulting in faster progress.
Deeper changes occur in the integration of the global development system.
The reporter learned that currently, Ascend has covered more than 90 mainstream third-party communities and has become the first Chinese computing platform that can be directly installed on PyTorch (one of the world's most mainstream AI development frameworks) official website.
Wang Tao also told reporters that previously, Chinese AI computing platforms needed to enter the PyTorch ecosystem, and developers needed to first go to Huawei's own community to download and install, which was a cumbersome process.
After months of effort, this obstacle has been eliminated, and global developers can now directly obtain Ascend's development tools on the PyTorch official website.
In the future, China's open-source big models will gradually shift towards using Ascend for native pre training, and after completing pre training, move towards inference and training of various vertical classes, without the need for additional adaptation.
System engineering is not just about hardware, but also about software ecology, involving operators, frameworks, and complex tool systems.
Kunpeng has over 4.16 million global developers and over 20 million installed openEuler systems, ranking first in China's server operating system market share.
The Lingqu protocol, along with reference designs and testing specifications, is fully open and donated to the Global Computing Alliance, with continuously updated versions, allowing enterprises without the ability to independently develop complex systems to enter the hypernode ecosystem.
The unification of internet standards avoids the duplication of investment and ecological fragmentation caused by individual efforts.
On the chip roadmap, the R&D progress of the Ascend 960 chip has exceeded expectations. From now on, Ascend chips will adhere to one generation per year, based on Lingqu interconnection, and super node clusters can support up to one million cards. Huatai Securities estimated in its research report in April this year that the market size of China's super node architecture is expected to reach 341.4 billion yuan by 2028.
A clearly defined and open-source computing power base is taking shape at the bottom of China's AI industry.
When the division of labor becomes mature, model companies no longer need to worry about base suppliers becoming competitors, application developers no longer need to duplicate wheel making, and computing power becomes a public supply for the industry like water and electricity.
On this silicon-based black soil, crops have already begun to grow, and with more and more industrial forces flowing in, even greater harvests are still ahead.