The New Pattern of AI Computing Power

Economic Observer Follow 2026-08-21 22:22

Xie Zuqian, Jiang Yiming/Wen

In our consulting work, more and more corporate clients are asking: What direction will computing power develop in the future? Today, when people talk about AI computing power, attention is usually focused on large data centers, high-end GPUs (graphics processing units), and increasingly large training clusters. But will the future still be like this?

As is well known, leading global technology companies are investing huge amounts of money in building new AI data centers. Meta's Hyperion campus in Louisiana has officially expanded its total investment to over $50 billion, with IT computing capacity increased to 5 gigawatts (GW). This project will become Meta's largest AI data center and one of the world's largest AI infrastructure projects.

To train larger models, these facilities require more chips, faster networks, stronger power supply, and more complex cooling systems. Large scale computing power parks seem to have become the most important infrastructure in the AI era.

However, will the development of AI computing power only move towards continuous concentration and large-scale?

We don't think so. While large-scale training centers continue to expand, another change is also happening. AI computing power is gradually entering decentralized regional data centers, enterprise data centers, factories, automobiles, robots, and other intelligent devices.

The future of AI computing power will not simply continue to move towards concentration or begin to disperse, but will form two coexisting trends: the organization of computing power pre trained by cutting-edge models will still be relatively concentrated, while inference will show more obvious differentiation according to task types.

But this does not mean that the importance of the cloud will decrease. Frontier and complex reasoning will still run extensively in the cloud, and enterprise privatization deployments will continue to increase, with more reasoning entering the edge and end side. Different computing tasks will run in different levels of infrastructure.

This change will not only reshape the landscape of AI infrastructure, generate new chip demands, but also further change the organization, processes, and decision-making methods of enterprises.

Why is cutting-edge pre training still relatively concentrated

Frontier model training is a highly complex system engineering.

Training large models typically requires thousands of GPUs or other AI processors to work together to complete the same task. Close, continuous, and low latency communication must be maintained between chips. Any communication bottleneck may reduce the efficiency of the entire cluster.

Therefore, the ability to train is not only determined by the performance of a single chip, but also by whether the entire system can effectively connect a large number of chips.

Nvidia improves the communication capability between GPUs through NVLink (GPU high-speed interconnect technology) and NVLink Switch (high-speed switching chip). Google's TPU (Tensor Processor) system also organizes thousands of chips into a unified computing node cluster through a dedicated high-speed network.

The common feature of these systems is to concentrate a large amount of computing power, allowing different chips to operate as a whole. This "concentration" mainly occurs in the pre training stage of cutting-edge models, as it highly relies on large-scale, high bandwidth, and low latency chip collaboration. Large AI data centers will therefore continue to undertake cutting-edge model pre training and other highly synchronized computing tasks.

But the relative concentration of computing power does not mean that model development will be further concentrated in the hands of a few companies. With the development of open source models, more and more enterprises and research teams can conduct secondary development on top of the basic model.

Some enterprises with deep professional knowledge and data barriers, such as those in the fields of finance, industry, and medicine, may further fine tune, continue training, and adapt their business to form model capabilities that are more closely aligned with their own business. The computational power required for these tasks is much smaller than pre training cutting-edge models and can be provided by cloud platforms, professional service providers, or enterprises themselves.

Most enterprises in the consumer goods, retail, and general service industries often do not require further model training, but rather utilize existing generic models and integrate their knowledge, data, and business tools with them.

But training is not the entirety of AI.

Reasoning is moving towards differentiation

Inference is the actual operation of a model after training, such as answering questions, running agents, recognizing images, analyzing on-site information, or assisting enterprises in making judgments.

Unlike cutting-edge training, many inference tasks can be completed independently without the need for thousands of chips to collaborate simultaneously. They can run in the cloud based on the location of users, data, and applications, or be deployed to internal enterprises, edge nodes, or edge devices.

In recent years, with the development of long-range agents, contextual engineering, and multi-step tool invocation, the computational requirements for some complex reasoning tasks have been increasing. They need to handle longer contexts, call more models and tools, and maintain a continuous running state. In the foreseeable future, most complex reasoning may still remain in the cloud. Cloud data centers can share computing power among different users and tasks, and reduce unit inference costs through high device utilization, mature software tools, and professional operations.

On the other hand, on-site perception, real-time prediction, and partial inference related to device operation have stronger distributed deployment conditions. As AI enters more practical scenarios, the two forces driving inference to extend beyond the cloud have become increasingly clear.

The first type of force is the privatization deployment within the enterprise.

Some data in fields such as healthcare, industry, finance, and government have high sensitivity and are subject to different regulatory requirements, making it difficult to easily access public clouds. The most important reasons for these fields driving AI from public clouds into their controllable environments are often data security, privacy, regulatory requirements, and control over data and systems.

In these scenarios, enterprises can run AI through their own data centers, private clouds, or dedicated computing nodes, keeping sensitive data and critical business within their controllable environment.

Microsoft's distributed infrastructure solution Azure Local allows enterprises to run cloud services and AI models on their local infrastructure. This allows some sensitive data to remain within the enterprise while completing inference and analysis.

The second type of force is edge and edge computing. Both bring computing power closer to where data and business occur, but in different locations. Edge computing mainly runs on servers or computing nodes near business sites, such as edge nodes in factories, parks or operator networks; End side computing runs directly on cars, robots, cameras, and other smart devices themselves.

The main driving force of edge and edge computing is that data generation and business execution are already on-site, and some tasks have high requirements for response speed, network continuity, and offline operation.

The distance itself will directly affect the response speed of AI systems. The data is transmitted from the site to the cloud, and then the results are returned, which requires multiple network transmissions and processing. The farther the node is, the higher the delay and the greater the uncertainty.

For general knowledge Q&A or backend analysis, a brief delay usually does not result in serious consequences. But in some critical scenarios, the requirement for real-time performance is very high. For example, car braking and emergency shutdown of industrial equipment require the system to respond in a very short period of time.

Therefore, in tasks involving safety and real-time control, critical control links must remain on site. The cloud can undertake relatively non real time tasks such as model training, version updates, and data analysis, while edge or edge AI can participate in perception, prediction, and auxiliary judgment, and operate in conjunction with validated real-time control systems.

In other words, tasks that are more suitable for edge and edge deployment are not necessarily complex large-scale model inference, but rather tasks that are closer to the physical world such as perception, on-site prediction, and auxiliary judgment.

The computing power required for a single end device may not be very large, but the number of cars, robots, and various intelligent devices can be very large. Therefore, in the long run, the end side may become one of the important sources of growth for distributed inference.

Enterprises such as Amazon Web Services (AWS) have begun deploying some real-time inference capabilities to factories, sites, and even devices, and supporting some nodes to continue running offline.

AI infrastructure is forming a multi-layered system

The future AI infrastructure may form a multi-layered system from the cloud to the edge, and then to the end side. Large AI data centers undertake cutting-edge model pre training and complex inference; Edge nodes undertake real-time inference and data processing near the business site; End side devices directly perform perception, prediction, and auxiliary judgment on devices such as cars and robots. The final control of safety critical tasks is still carried out by a real-time control system that has been validated on site.

At the same time, for enterprises with high requirements for data security, privacy, and compliance, some AI will also run privately in their own data centers, private clouds, or edge nodes.

This multi-layered system is also affected by energy conditions. The International Energy Agency predicts that by 2030, global data center electricity consumption may nearly double, reaching approximately 945 terawatt hours (TWh).

Once a large computing power center is built, its geographical layout is usually relatively fixed, and power supply becomes the main constraint affecting its continuous expansion. The power supply capacity, transmission and distribution conditions, and grid connection level will directly affect the expansion space and upper limit of the computing power center.

In contrast, the location of edge nodes depends more on the business site, real-time requirements, and network conditions. The end side computing power on cars, robots, and other intelligent devices must follow the device itself, making it difficult to migrate solely based on energy costs.

Therefore, rather than saying that distributed computing is "solving" energy problems, it is more accurate to say that it is reshaping the spatial distribution logic of computing power: large-scale training and complex inference will still mainly run in computing power centers with stable power and infrastructure conditions, while more on-site inference will enter the edge and end side as business needs arise. At the same time, some enterprises may also adopt private deployment due to data security and control requirements.

The new computing power layout will give birth to new chips

Distributed inference does not mean moving large training GPUs directly into small computer rooms or smart devices.

Training GPUs aims to achieve the highest possible computing power, large-scale parallel capability, and high-speed interconnection. However, chips deployed in edge nodes and end devices are often limited by power, space, heat dissipation, and cost.

These scenarios require high energy efficiency, low latency, larger local memory, and more compact and easily deployable hardware.

Even if the basic structure of future AI models gradually converges, there is still a lot of customization space for inference chips. The requirements for model size, computational accuracy, memory bandwidth, response speed, power consumption, and cost vary in different scenarios.

The advantage of a general-purpose GPU is its flexibility, which can support different models and tasks. But this universality also comes at a cost in terms of efficiency. In inference scenarios with relatively stable workloads and clear task boundaries, chips designed for specific models, computing methods, or device environments can often achieve higher energy efficiency and lower operating costs.

This does not mean that general-purpose GPUs will be replaced, but rather that the AI chip market will form a clearer division of labor. Large GPUs continue to undertake model training and some complex inference, while cloud inference chips serve large-scale workloads. Edge chips support inference close to the business site, while edge chips are more responsible for real-time perception, prediction, and auxiliary judgment of the device itself.

GPUs, NPUs (neural network processors), and various specialized accelerators will serve different workloads as a result, and new architectures such as storage computing integration may also be developed in some inference scenarios.

Recently, we have explored the trend in this area with Rear Mobility Intelligence. Its Marvel M50 is designed for large-scale model inference on both the end and edge sides, using an integrated storage and computation architecture to reduce the repetitive transfer of model parameters between storage and computing units, thereby improving inference efficiency.

This type of chip may not completely replace large GPUs, but it is likely to further expand the boundaries of the AI semiconductor market.

In the future, competition among enterprises will not only lie in who has the strongest training chips, but also in who can find a better balance between performance, cost, power consumption, and deployment difficulty for different inference scenarios.

Of course, whether new chips and computing architectures can truly form commercial value ultimately still needs to enter real-world scenarios and undergo various tests such as performance, cost, reliability, and deployment difficulty. In this regard, China may also become an important testing ground.

On the one hand, China has a large number of manufacturing, automotive, robotics, logistics, energy, and urban infrastructure scenarios. These industries have a significant demand for AI in real-time response, local processing, and cost control.

On the other hand, China also has a complete industrial system from chips, servers, operator networks, and industrial parks, to complete machine manufacturers, automobile companies, robot companies, and a large number of industrial users. This enables new chips, models, and solutions to quickly enter real business environments, be continuously validated in applications, and be continuously adjusted in collaboration across the upstream and downstream of the industry chain.

The policy level is also driving this trend. In January 2026, the Ministry of Industry and Information Technology and other eight departments released the implementation opinions of the special action of "artificial intelligence+manufacturing". While promoting the development of training chips, they also proposed to develop end-to-end reasoning chips, edge computing servers, and promote the construction of the "cloud edge end" model system.

In such an environment, what enterprises need to verify is not just the performance of a single chip, but whether the chip, model, software, network, data, and business processes can be combined into a complete solution.

A new technology can first enter a few scenarios, be repeatedly adjusted in practical applications, and then be promoted to more scenarios through operators, whole machine enterprises, and industry partners, forming a cycle of "experiment adjustment scale".

Therefore, China may not only become a testing ground for new AI chips, but also an important testing ground for the large-scale application of new AI infrastructure and solutions.

New computing power grid reshapes enterprise organization

The changes in computing power structure have profound implications for enterprises.

Many companies today have both headquarters personnel, regional teams, and frontline employees. The headquarters is responsible for strategy, standards, resource allocation, and shared capabilities; Frontline employees are closer to customers, equipment, and actual business, and need to understand the on-site situation more quickly and respond accordingly.

As AI enters more and more business scenarios, enterprises no longer only need to manage backend models and tools. Today, many intelligent agents are based on a universal large model, connecting enterprise knowledge bases, software tools, and execution interfaces. When these intelligent agents are able to invoke enterprise knowledge and tools with clear responsibilities, certain permissions, and continuously complete specific tasks, they will become closer to the "digital employees" in the enterprise.

As digital employees gradually increase, they may also form a division of labor similar to human organizations. A portion of digital employees operate on a central platform, mastering common knowledge, rules, and capabilities of the enterprise; The other part operates in specific business scenarios such as regions, factories, and stores, calling on local data and systems to undertake tasks closer to the front line.

The way enterprises manage digital employees will become increasingly similar to that of managers. Enterprises need to decide what data they can access, what knowledge they have, what permissions they have, and which tasks they can independently complete.

From this perspective, the ultimate management of enterprises is not just AI, models, or computing power, but a new type of organization composed of human employees and digital employees.

This change first requires enterprises to change their mindset of 'one deployment method to solve all problems'. Complex reasoning can continue to leverage the scale advantage of the cloud; Sensitive data and regulated businesses can be deployed privately; Tasks that require proximity to the business site can enter the edge; In scenarios where data needs to be processed directly on the device, end-to-end computing can be further adopted.

The deployment methods are different, but the underlying principle is the same: different AI applications should be placed in the most suitable location according to actual business needs.

This approach also provides a more realistic path for the AI transformation of large enterprises. Enterprises do not need to push for a large-scale transformation that covers the entire company from the beginning. They can start with a position, a team, a production line, or a specific business problem, and form a closed loop that can operate on a smaller scale.

BMW Group's Factory Genius provides a specific example. It is an AI assistant designed for factory equipment maintenance. When production equipment malfunctions, maintenance personnel can directly ask it questions. The system will search for relevant information from equipment manuals, quality data, internal fault reports, planning documents, and daily updated shift records, and quickly provide specific suggestions for the problem, thereby helping maintenance personnel locate and handle the fault faster.

Factory Genius was initially piloted at the Dingolfin factory in Germany, and similar explorations were also conducted at factories such as Spartanburg in the United States and Rosling in South Africa. Subsequently, the AI team at BMW Munich headquarters integrated the needs and experiences of different factories to form a group level application, and promoted it to more factories through internal platforms.

BMW's approach shows that AI applications can be explored from specific business problems and scenarios, and after verification, the headquarters can integrate experience, accumulate common capabilities, and gradually promote them to other business units.

Of course, starting small does not mean going our separate ways.

In the future, it is highly likely that the AI infrastructure of enterprises will not be provided by a single cloud platform, hardware system, or vendor, while using public and private environments, and deploying edge nodes and edge devices in some businesses. Different models and chips may also come from multiple suppliers.

But using multiple suppliers does not mean that companies should infinitely increase their technology stack. The headquarters needs to establish a common base, standards, and rules, while business units can conduct experiments and applications according to actual needs. Enterprises need to retain the right of choice and system resilience, while also avoiding the establishment of separate AI systems by different business units.

Enterprises should pursue "centralized governance, distributed application, and controlled technological system", rather than relying on the central government to remedy the situation after the technology is infinitely dispersed.

Digital employees also need to be trained, evaluated, and continuously updated like human employees. Enterprises need to clarify their knowledge sources, data and authority boundaries, division of responsibilities with human employees, and how to continuously correct behavior based on actual work results, and transform the experience formed on the front line into a common capability for the entire organization.

The deployment of AI infrastructure will affect where digital employees operate and how they access data and systems in different business scenarios. The entry of digital employees into different business scenarios will further affect how information flows, where decisions occur, and how human employees and digital employees divide labor.

The technical architecture, organizational structure, and business processes will gradually merge into one system. With the continuous accumulation of digital employees' abilities, their mastery of enterprise specific knowledge, memory, and collaboration methods may gradually become a new organizational asset.

In this sense, AI is not only changing the way businesses operate, but also reshaping the value logic of enterprises.

AI competition is entering a new stage

AI competition depends not only on who has the largest training cluster.

Large data centers will continue to push the boundaries of cutting-edge model training and complex inference capabilities; At the same time, different inference tasks will run in the cloud, edge nodes, or end devices based on data, business scenarios, and real-time requirements; For enterprises with high requirements for data security and control, some AI will also adopt private deployment. Correspondingly, different chips will also form clearer division of labor.

Deeper changes occur within the enterprise. As digital employees enter more and more business scenarios, companies need to rethink the division of labor between them and human employees.

The real challenge is not to purchase more chips, but to organize computing power, data, digital employees, and human employees into a whole that can continuously learn, collaborate effectively, and operate at scale.

Enterprises not only need AI technology, but also a new system thinking. What we are witnessing is not only a new paradigm of AI computing power, but also a new era of organizational science.

(Xie Zuqian is the founder and chairman of Gaofeng Consulting Company, and Jiang Yiming is a partner and general manager of Gaofeng Shuzhi.)