
In 2025, there will be a significant change in the global AI industry.
The focus of competition for large models is shifting from parameter scale to deep penetration into application scenarios. The computational power required to train a large model is still growing exponentially, but at the same time, embodied intelligence AI Agent、 The requirements for chips in new directions such as digital twins are also rapidly differentiating.
Training a large model requires high-density AI computing, training a strategy model for a robot requires running a physics engine and graphics rendering simultaneously in a simulation environment, and deploying an end side agent requires multitasking parallel processing under low-power conditions.
In the past, the common practice in the AI chip industry was to design dedicated chips for each type of task. However, as customers increasingly need to handle these tasks simultaneously in the system, whether a single chip can cover AI computing, graphics rendering, physical simulation, ultra high definition video encoding and decoding, and scientific computing has become an unavoidable issue.
This is the background of the attention paid to the technology route of full function GPU. Full function GPU refers to a GPU that can simultaneously support these five types of tasks on a single chip architecture.
Renowned market research firm Frost&Sullivan predicts that the Chinese AI chip market will increase from 142.5 billion yuan in 2024 to 1.3 trillion yuan in 2029, with an average annual compound growth rate of 54%.
Recently, Moore Thread (688795. SH), a leading domestic enterprise in the field of fully functional GPUs, released its semi annual report for 2026.
Moore Thread was established in 2020 and has successively launched five generations of GPU chip architectures including Sudi, Chunxiao, Quyuan, Pinghu, and Huagang, building a complete product matrix covering cloud, edge, and terminal. It will be listed on the Science and Technology Innovation Board in December 2025.
As of the first half of this year, the company has a research and development team of 1019 people, accounting for 74.16% of the total number of employees, of which over 70% of the R&D personnel have a master's degree or above.
In the first half of this year, Moore Thread achieved a revenue of 1.736 billion yuan, a year-on-year increase of 147.42%. The revenue scale in the first half of the year has exceeded that of the whole year of 2025, with a total gross profit of 989 million yuan and a gross profit margin of 56.95%. The net profit attributable to the parent company has significantly narrowed compared to the same period last year.
On the same day, Moore Thread announced the launch of H-share issuance and preparation for listing on the main board of the Hong Kong Stock Exchange.
As AI moves from language to the physical world, the core contradiction of computing infrastructure also changes. In the past, we were comparing who had a higher peak AI computing power, and now we are comparing who can provide comprehensive training, rendering, and simulation capabilities on a single chip.
At this turning point, the value of fully functional GPUs is being re recognized.
Comprehensive breakthrough of cluster business and accelerated commercialization
Behind the performance growth of Moore's Thread in the first half of the year, the increase in cluster business volume is the most direct driving force.
The company's flagship product MTT S5000 belongs to the fourth generation "Pinghu" architecture, with a single card dense AI computing power of 1000 TFLOPS (10 quadrillion floating-point operations per second), equipped with 80GB of video memory, a video memory bandwidth of 1.6TB/s, and a card to card interconnect bandwidth of 784GB/s. It fully supports full precision computing from FP8 to FP64. A S5000 can handle both large model training and inference tasks simultaneously, which is the core feature that distinguishes it from dedicated AI acceleration chips.
The KUAE intelligent computing cluster based on S5000 achieved large-scale sales in the first half of the year and was deployed and delivered in multiple cities such as Beijing, Wuxi, and Hangzhou. In its semi annual report, Moore Thread positioned itself as one of the few GPU suppliers in the market that truly realizes the commercialization of large-scale cluster applications at the kilocard and ten thousand card levels under a single network for training.
For customers who purchase large-scale computing power, whether a cluster can operate stably in a long-term, high load environment is more critical than single card scoring.
The Kuai'e cluster has quantifiable results in this dimension, with a linear scaling efficiency of 95% for cluster training. This means that as the cluster size expands from kilocalories to tens of thousands of calories, the loss of computational efficiency is controlled within 5%.
The computational power utilization (MFU) of the Dense model reaches 60%, while that of the MoE model reaches 40%. The cluster supports breakpoint continuation training, which can quickly recover in case of hardware failure causing training interruption, and the effective training time accounts for over 90%. Under the native FP8 precision, the cluster has fully reproduced the training process of top large models, and multiple key indicators have reached the international mainstream level.
It is worth noting that the S5000 (PH100 chip) has also passed the national "Security and Reliability Assessment" for the first time, opening up a compliance access channel for Moore Thread in industries with extremely high information security requirements such as government, finance, and energy. During the reporting period, Moore Thread was also included in the Sci Tech Innovation 50 Index.
The large-scale sales of clusters have solved the problem of whether they can be sold, but the market's deeper doubts about domestic GPUs lie in their training capabilities. The saying that 'domestic GPUs can only perform inference and cannot support high-end training' was once the most widely circulated in the industry. The set of practical results submitted by Moore Thread in the first half of the year responded positively to this issue.
Based on the Kuai'e cluster, our partners have trained a MoE-236B basic large model with over 25 trillion tokens of corpus from scratch. The EvoPhys team at Peking University trained the world's first 5D world model EvoPhys World on the S5000, which won first place on the "World Generation" track publicly evaluated by Stanford University's WorldScore and remained at the top for 37 consecutive days.
Moore Thread, in collaboration with Beijing Zhiyuan Artificial Intelligence Research Institute, completed the full process training of the embodied brain model RoboBrain2.5 based on the FlagOS Robo framework. The company's open-source code model MusaCder has kernel operator generation performance comparable to international SOTA levels.
From big language models to world models and embodied intelligence models, this set of training results covers several cutting-edge directions in the current AI industry, which also means that the capabilities of domestic GPUs have expanded from inference to high-end training.
Training ability is the threshold, and ecological compatibility determines whether customers are willing to cross this threshold.
The MUSA software stack developed by Moore Thread has achieved 100% compatibility with core mathematical libraries and full compatibility with over 3000 PyTorch operators, covering 55 types of core AI operators. The world's top inference frameworks vLLM and SGLang have both provided official support for MUSA. SGLang has completed the mainline integration of MUSA code, and the open-source ecosystem of domestic GPUs has entered the native support stage.
In terms of large-scale model adaptation, Moore's Thread has implemented Day-0 adaptation for DeepSeeker V4, MiniMax M3, Zhipu GLM-5.2, as well as the trillion parameter model Kimi K3 and multimodal model MiniMax H3, which means that the model was validated on domestic GPUs on the day of its release. As of the end of the first half of the year, the number of developers in the MUSA ecosystem has exceeded 800000.
Zheng Weimin, an academician of the CAE Member and a professor of Tsinghua University, made a judgment at the World Conference on Artificial Intelligence that the core indicators for measuring computing power infrastructure have shifted from peak computing power values to the number and quality of effective words that can be stably produced per unit computing power.
This judgment actually points to the value of fully functional GPUs in the computing power industry. The more complete the abilities of training, reasoning, rendering, and simulation, the higher the efficiency of converting unit computing power into effective output.
In terms of customer structure, Moore thread is accelerating its penetration into two types of core computing customers: the Internet and operators. On WAIC 2026, Moore Thread released the results of cooperation with many leading enterprises in the industry, such as Best Vision and Frontier Technology, which demonstrated the company's ability to implement cutting-edge technology scenarios such as big model training and reasoning, embodied intelligence, and world models, and output mature domestic computing solutions for key industries such as energy, finance, education, and the Internet.
The market generally believes that Internet and operator customers have large procurement scale, long technical certification cycle, and strong continuity and stability of orders once they enter the supply chain.
In addition, Moore's Thread's product line is not limited to cloud computing. On the edge side, the MTT E300 AI computing module based on the self-developed "Yangtze River" SoC chip provides localized AI computing power for industrial quality inspection, energy inspection, embodied intelligence, intelligent vehicles, and low altitude economy scenarios.
On the terminal side, the company has launched the AI computing laptop MTT AIBOOK and the home AI hub MTT AICUBE, the latter of which has been launched for sale on the JD platform. From cloud computing to edge inference and then to terminal applications, Moore's Thread's products cover the complete chain of computing power demand from centralized to decentralized.
In fact, the quality of Moore's Thread's narrowing losses in the first half of the year is also worth paying attention to.
GPU chips have typical high fixed cost and low marginal cost attributes. After the revenue scale expands, the initial R&D investment and chip production costs are diluted in larger shipments, and the improvement of the profit model is sustainable. A more crucial signal is that the company's R&D investment in the first half of the year reached 769 million yuan, a year-on-year increase of 38.16%. Since 2022, the cumulative R&D investment has approached 5.9 billion yuan.
The narrowing of losses is not a short-term improvement obtained by compressing long-term investments, but a result of improving commercial efficiency. The pace of technological iteration has not slowed down due to the pursuit of profits.
Produce the next generation of chips and rush to the global market
The direction of R&D investment growth for Moore Thread in the first half of the year was very concentrated. In its semi annual report, it clearly stated that the main reason for the year-on-year increase in R&D expenses was the focus on investing in two new chips, Huashan and Lushan, under the Huagang architecture.
Public data shows that the current investment in the development of Moore Thread Huashan chips and applications is 275.8373 million yuan, with a cumulative investment of 605.8679 million yuan. The current investment in Lushan chip and application research and development is 197.4846 million yuan, with a cumulative investment of 336.08 million yuan.
The total investment of the two projects in this period is about 473 million yuan, accounting for about 60% of the total R&D investment in the first half of the year. Both chips are in the research and development stage, and their respective investment amounts in this period have approached or even exceeded the cumulative investment of all previous periods. R&D is entering an intensive period.
Huagang is the fifth generation fully functional GPU architecture of Moore's Threads, released in December 2025, supporting full precision computing from FP4 to FP64, with a 50% increase in computing power density and a 10 fold increase in energy efficiency compared to the previous generation. It can support cluster expansion of over 100000 cards. Based on the Flower Harbor architecture, Moore Thread has planned two product lines.
Huashan's integrated training and promotion scenario for data centers, the company stated in its semi annual report that it will have multiple leading or even surpassing international mainstream chip capabilities in floating-point computing power, memory access bandwidth, memory access capacity, and high-speed interconnection bandwidth.
Lushan focuses on high-performance graphics rendering, while also possessing the ability to infer large models. The rendering performance of 3A games has improved by 15 times compared to the previous generation, AI performance has improved by 64 times, and ray tracing performance has improved by 50 times.
These two chips respectively point to the two largest directions in the current AI chip industry, cloud computing and high-end consumer market. Large model training and large-scale reasoning are currently the fastest growing and highest gross profit segments of the global semiconductor industry. The budget of the leading Internet companies and telecom operators is still expanding rapidly.
The high-end consumer graphics card market has long been dominated by overseas giants, and domestic GPUs have not yet produced truly competitive products in this direction. If Lushan continues to advance as planned, Moore's Thread has the opportunity to become the first domestic manufacturer to enter the high-end consumer graphics card market.
Based on the "full functionality" advantage of its own product, Moore Thread has launched the MT Lambda full stack simulation platform in the field of embodied intelligence. It has open sourced MuJoCo Warp MUSA, the first GPU accelerated physical simulation backend based on MUSA architecture in the embodied intelligence field, and completed Sim2Real real machine verification of quadruped robotic dogs and bipedal humanoid robots. At the same time, Moore Thread has established an industrial embodied intelligent innovation center in Wuxi, jointly building with the first batch of 16 ecological partners. It has also reached a strategic cooperation with Xiaoma Zhixing to accelerate the large-scale landing of L4 level autonomous driving with domestic computing power.
From cloud based large-scale model training to strategy learning in simulation environments, and then to end-to-end deployment on real machines, Moore's Thread relies on the comprehensive capabilities of fully functional GPUs in AI computing, graphics rendering, and physical simulation to build a complete technical link from training to implementation in the field of embodied intelligence.
From the demand side, the growth certainty in the above directions is strong.
In August 2025, the State Council issued the "Opinions on Deepening the Implementation of the" Artificial Intelligence+"Action", which included "strengthening the overall planning of intelligent computing power" as one of the eight basic supporting capabilities.
Frost&Sullivan predicts that China's intelligent computing power will grow from 59.20 EFLOPs in 2020 to 438.07 EFLOPs in 2024, with an average annual compound growth rate of 64.9%. It is expected to continue to grow at a compound growth rate of 45.3% from 2025 to 2029, reaching 3035.91 EFLOPs.
While demand continues to increase, the validation of domestic GPUs in terms of training capabilities and ecological compatibility is constantly accumulating, and the market share that can be undertaken is gradually expanding.
On the same day, Moore Thread announced the start of preparations for the H-share issuance. In the first half of 2026, 24 A-share companies have completed H-share issuance, raising a total of approximately HKD 121.7 billion. The collective construction of cross-border capital platforms by hard technology companies has become one of the clearest trends in the capital market this year.
The Moore thread chooses to initiate H-share preparation at the node where the current generation of products is in high volume and the next generation of chips is being reinvested. On the one hand, we have seen clear signals from the policy side, supporting leading enterprises to make good use of "two markets and two resources". In August 2026, the China Securities Regulatory Commission publicly stated that it actively supports mainland enterprises that meet the conditions to list in Hong Kong; On the other hand, it is also a lever for opening up a global layout. The international brand reputation and overseas channel resources brought by the listing of Hong Kong stocks will become an important driving force for the next generation of chips to enter the global market, helping the company truly open up the second growth curve of performance.
At the World Artificial Intelligence Conference in July, Zhang Jianzhong, founder, chairman, and CEO of Moore's Thread, once said, "You can rest assured that model training on large-scale domestic clusters has been done very well today. Moreover, Moore's Thread can provide localized services to users at any time, which is definitely better than foreign companies