Economic Observer Follow
2026-08-23 08:56

Economic Observer reporter Zheng Chenye
From August 19th to 23rd, the 2026 World Robot Conference was held in Yizhuang, Beijing. More than 300 companies showcased over 2000 exhibits and launched over 150 new products.
Compared to a year ago, the appearance of the conference has undergone significant changes: last year, most of the robots on the exhibition booths were still performing performance actions such as somersaults and boxing, but this year they have switched to practical displays such as sorting drugs in simulated pharmacies, picking up goods for customers in retail scenes, and folding clothes and organizing items in simulated homes.
But the reporter observed on site that robots also occasionally encounter problems in practical scenarios.
On the afternoon of August 20th, the reporter noticed a long queue in front of an unmanned retail booth, with many onlookers. A set of shelves has been set up on the booth, containing both medicines and daily necessities. Next to it is a checkout counter, where customers can place orders through a mobile mini program. After receiving the order, the robot retrieves the goods from the shelves and places them at the checkout counter to complete the delivery.
A spectator placed an order for a bag of soybean milk. After receiving the order, the robot first turns to the shelf to identify the position of soybean milk, reaches out the mechanical arm to grab it, and turns to put it on the checkout table. The whole movement is smooth and coherent. More and more people were watching, and another viewer placed an order of small items on the inside of the shelf. But this time, there was a noticeable pause in the robot's movements. The staff explained that this is because there is too much pedestrian traffic nearby, which interferes with the robot's obstacle avoidance recognition. Some viewers said that with such a small flow of people, it would get stuck when picking up two items. If it were to work in a retail store, there would definitely be more interference than here, and it would definitely not be able to do it.
In fact, the above situation is related to the generalization ability of robots: humanoid robots currently do not have a brain that can flexibly respond to various situations like humans; The behavior of robots relies on data collection and scene simulation training in the early stage, essentially executing according to preset programs. Trained actions can be completed, but once untrained situations occur, it does not know what to do.
Customers buy robots to solve problems, not to serve them, "said Duan Yanbiao, CPO (Chief Product Officer) of Yunji Technology, when talking about the challenges encountered by robots in scene implementation.
Yunji has been deeply involved in the field of service robots for many years, deploying service robots in over 40000 hotels worldwide. Duan Yanbiao said that the reason why customers "serve" robots in reverse is because robots encounter a large number of long tail problems during operation (i.e. unexpected situations with low probability of occurrence but diverse types). Taking the hotel scene as an example, the long hair on the carpet sometimes sinks the wheels of the robot, which is difficult to cover in advance in the laboratory and simulation training. Once it occurs, customers need to spend time dealing with it specifically.
Duan Yanbiao believes that these long tail issues are one of the main reasons hindering the large-scale commercialization of robots.
On August 20th, Wang Xingxing, founder of Yushu Technology (688836. SH), said at the main forum of the conference that the biggest bottleneck in the world is the insufficient generalization ability of embodied intelligence. Currently, most models can achieve a task success rate of nearly 100% after sufficient data collection and specialized training in fixed scenarios, but as long as the operating items are changed or the environment is slightly changed, the success rate will significantly decrease.
In Wang Xingxing's view, when a robot can complete about 80% of tasks in any unfamiliar environment, it is almost the ChatGPT moment of embodied intelligence (i.e. the breakthrough of intergenerational capabilities similar to ChatGPT in embodied intelligence), which can take as fast as two to three years, and as slow as five to ten years.
In this context, a competition has been launched around the accumulation of data, model training, and exploration of business models for robot "brains".
Laying eggs along the way
On a fixed workstation with high certainty, the success rate of robot operations is only slightly different from that of manual labor; Once the task becomes longer and more complex, this gap will quickly widen.
Xiaomi's new generation humanoid robot "Tieda" has been working on the nut workstation at Xiaomi's car factory since March this year. The nut workstation is a fixed workstation on the automobile assembly line. Workers need to align nuts with screw holes one by one on the assembly line, tighten and fix them. Each round of operation lasts for a few seconds and is repeated hundreds of times a day. The action is highly repetitive, but there are requirements for positioning accuracy.
The general manager of Xiaomi's robotics division, Xiang Diyun, told reporters that the success rate of "Tieda" when it first arrived at the workstation was 90%; After 4 months, the success rate has increased to 98%; The success rate of manually completing the same operation is 99%, and this last percentage point may catch up by the end of the year.
On fixed workstations with high certainty such as nut workstations, the efficiency of robots is highly correlated with task duration. Xiang Diyun said that in a fixed workstation, the efficiency of robots can continuously approach that of manual labor, but once the task becomes longer, more complex, or there are variables in the environment that are not covered by training, the success rate will quickly decrease. For example, for short tasks of 3 to 5 seconds, the efficiency of robots can reach 70% to 10% of that of humans; For medium tasks of 30 to 60 seconds, the efficiency of robots drops to 60-70% of that of manual labor; Long tasks lasting more than one minute can reduce efficiency to about 30% of manual labor.
Wang Xingxing explained in the above speech the technical reasons for this efficiency decline: every perception and control loop of the robot will produce deviation and loss, and the longer the task, the greater the accumulated deviation.
An investor who has long been interested in the field of robotics told reporters that in the robot race, "99 points equals 0 points", and what does a difference of 1 point mean to users? Just look at the robotic vacuum cleaner. Robotic vacuum cleaners have been in ordinary households for many years, with highly vertical tasks and relatively limited functions. Nevertheless, users still need to make a round of preparation before each startup: any scattered debris on the ground should be collected first because the machine cannot recognize it; The doors of some rooms need to be closed first because the machines cannot cross the threshold.
The investor believes that users buy robots to have the machines do their work for them. Once the machines cannot handle slightly complex situations, users have to take care of the machines in turn. The more complex the tasks and the more open the scenarios, the higher the cost.
If the robotic vacuum cleaner is still like this, the problem will only be bigger for general humanoid robots.
At this year's World Robot Conference, the reporter noticed that the home is the scene where companies invest the most in showcasing various household tasks, from folding clothes and cooking, organizing items, to restocking refrigerators. But all manufacturers' home scene demonstrations are still in the prototype display stage. No product from any company can be officially put into the homes of ordinary people.
Founder of Galaxy General Motors, Wang He, said on the main forum of the conference that, given current technological conditions, directly introducing robots into ordinary households is not a very good development path. Mo Lei, Vice President of Zhifang, also believes that it will take at least 5 years for humanoid robots to enter the household scene.
Additionally, being able to complete a task and being able to complete a task well are two different things. At the conference, reporters observed that some companies' robots take about 10 minutes to fold clothes once. In a robot dining table cleaning demonstration at a certain booth, a crumpled tissue stuck to the mechanical claw. The robot tried several times but did not throw it into the bucket. Finally, the staff stepped forward to handle it. During the exhibition, there were even cases where multiple robots were unable to continue demonstrating due to on-site network interference.
The family scenario is not feasible in the short term, and the current consensus in the industry is to transition step by step from the B-end.
The path proposed by Galaxy General is to first operate continuously in retail pharmacies, industrial production lines, and other scenarios, then transition to institutional health care and hospital wards, and then penetrate into essential family scenarios such as elderly care and disability assistance, ultimately spreading to ordinary households. Wang He called this process' laying eggs along the way '.
In the process of moving towards home, the abilities of robot brains need to continue to evolve.
At the technical architecture level, the robotics industry has experienced a directional convergence in the past year: by 2025, most embodied intelligence enterprises will adopt the VLA model (i.e. visual language action model, where robots simultaneously process the images they see, the language instructions they receive, and the actions they output); By 2026, most companies will begin to overlay world models on top of VLA, allowing robots to predict the future state of the environment before performing actions, reducing the accumulation of bias during task execution.
The industry has simultaneously formed an engineering consensus of "fast slow brain": the brain is responsible for deep reasoning and task planning, while the cerebellum is responsible for high-frequency real-time execution and motion control. The division of labor between the two balances computing power consumption and response speed.
Architecture is converging, but ultimately the brain needs data to become smarter. Currently, the robotics industry mainly relies on two paths to accumulate training data.
One is to deploy robots at the customer site to collect scene data. For example, Zheng Xiaodan, the head of JD Retail's embodied intelligence business, told reporters that robots are moving from "display" to "work". As a user of logistics, retail, health and other scenarios, JD hopes to open up its own scenarios and work together with robot manufacturers and model manufacturers to promote data collection. The goal for the next two years is to collect high-quality scenario data at the level of millions of hours.
Zheng Xiaodan said that measuring data cannot only focus on quantity, but more importantly, on high quality and diversity of scenarios.
The second is simulation. The principle of simulation is to build scenes in a virtual environment for robots to train repeatedly. The cost of virtual environment is much lower than deploying robots on customer sites, and it can quickly create various scene changes, such as changing lighting, moving object positions, simulating crowd interference. In theory, it can cover a large number of training scenes at low cost.
But the problem with simulation is that there is a gap between virtual environments and everyday scenarios. The physical laws in virtual environments are simplified, such as a tissue being a standard flexible object in simulation; However, it is difficult to simulate the tactile sensation of crumpled tissues sticking to mechanical claws and the trajectory of tissues that become heavier after being soaked in water slipping off the claw surface in daily scenarios. Simulation can stack up the amount of data, but it is difficult to cover the occasional, random, and unpredictable variables in daily scenarios.
In the industry, this gap is referred to as the sim to realgap (i.e. the transfer loss from simulation to reality), and robots trained solely through simulation often experience performance degradation after entering daily scenarios.
In Mo Lei's view, "working" is the highest ceiling scenario, and robots can only continuously expose problems, accumulate data, and iterate models by continuously working at customer sites. He believes that a physical intelligence company should not only make models, but should integrate software, hardware, and scenarios; The collaboration and iteration between the brain and the ontology are necessary. If we only focus on modeling without hardware or scene adaptation, the model's capabilities cannot be continuously improved.
In terms of cost and hardware, Mo Lei told reporters that if humanoid robots want to enter ordinary households, the hardware cost must at least reach the price range of A-class cars (about 100000 to 200000 yuan). At the same time, a robot can work continuously for 8 to 16 hours a day and operate without failure for 30 days a month. The importance of hardware stability may not be inferior to that of the brain itself.
As a reference, the current operating life of Tesla Optimus robot's dexterous hand is only 6 weeks, and the cost of a single hand exceeds $6000 (about RMB 42000). The annual cost of replacing only parts for a robot is nearly $100000 (about RMB 700000).
At the exhibition site, the vice president of a body intelligence company summarized to reporters the current position of the robotics industry in three sentences: data is the textbook, evaluation is the exam, and deployment is the job. At present, most robot companies are still in the stage of using textbooks for learning, and cannot guarantee 100% passing the exam, which is further away from large-scale employment in the future.
Chen Feng, the person in charge of Zhuji Power, said in an interview with reporters that 2026 is the first year of the landing of the POC (proof of concept) for humanoid robots, not the first year of mass production, and large-scale production will at least last until next year.
Chen Feng said that in the past, many robots performed at exhibition booths with an operator holding a remote control at the back; The smooth movements seen by the audience do not solely come from the robot's own judgment. The most important step for the humanoid robot to enter the scene is to "remove the remote control".
The Brain Business
The solutions to challenges such as insufficient generalization ability, inability to cover long tail problems, and inability to navigate household scenarios all depend on whether robots can have a sufficiently intelligent brain.
A new business is emerging around the brain.
Xinghai Tu founder and CEO Gao Jiyang said at the main forum of the conference that the business model of embodied intelligence is changing and will be divided into three stages in the future: the current stage is mainly focused on whole machine sales, with hardware gross profit margins maintained between 40% and 60%; The second stage is the subscription phase of the solution, where customers no longer purchase the entire machine but pay for the overall solution, and the hardware gross profit margin will fall back to around 20%; The third stage is the physical world token (i.e. the unit of measurement for the use of intelligent capabilities) sales stage, where customers pay based on the amount of intelligent capabilities consumed by robots to complete tasks, and the entire machine may even enter negative gross profit. The source of income shifts to the continuous consumption of intelligent capabilities themselves.
Gao Jiyang believes that this evolutionary path is similar to the process of the cloud computing industry from selling servers to selling computing power. The brain will gradually evolve from a component of the machine to an independent source of income, and hardware will in turn become the carrier of the brain.
The premise for the brain to become an independent business is that the brain itself is intelligent enough. The brain needs two things to become smarter - data and computing power.
An analyst from a large securities firm in southern China told reporters that the current problem in the robotics industry lies in the model, and the problem with the model lies in the data. If there are no large-scale companies in the data stage, it will be very difficult for stocks in the humanoid robotics sector to experience a large-scale upward trend.
Zheng Xiaodan said that data collection is a complete chain from front-end collection, cleaning and labeling, to evaluation. The collection methods include collecting data in the scene through robot bodies, collecting data through first person perspective devices, and other paths. JD.com is also expanding data collection in more industrial and home scenarios with external partners, but the standardization and scaling of the entire chain are still in the early stages.
In terms of computing power, Huafu Securities recently released a research report stating that Tesla completed the chip fabrication of its next generation AI chip AI5 in April 2026, with a single chip computing power of 2000 to 2500 TOPS (trillions of operations per second), which is 4 to 8 times that of the previous generation AI4. AI5 is the computing power foundation of Optimus' next-generation brain, directly determining how complex inference models the robot can run on the end side (i.e. the robot itself rather than the cloud).
Besides data and computing power, Tesla has an advantage that most competitors do not possess. Public information shows that Tesla FSD (Full auto drive system) has been included in the list of available regions in China in May and is conducting internal tests in Beijing, Shanghai and other cities. FSD and Optimus share end-to-end neural network architecture and visual perception algorithms. Tesla's automated driving data and algorithms accumulated on the road can be directly used to train robot brains.
The above research report also mentioned that Optimus is expected to achieve a weekly production of 2000 units by the end of the year, with a full year shipment target of 10000 units in 2026 and an estimated shipment of over 100000 units in 2027. Deploying more Optimus on customer sites will generate more scenario data to feed back into brain training.
In addition, more and more industry giants are entering the field of embodied intelligence.
The aforementioned investors stated that in the past few years, the technology breakthroughs of startups have supported the field of humanoid robots. However, starting from the second half of this year, the resource investment and scene reserves of industry giants have supported this field. Companies such as NVIDIA, Google, JD.com, and Huawei have entered from different directions, and their common entry point is the brain.
The aforementioned investor believes that at present, there is a lack of pure brain and data chain targets in the secondary market. Most listed companies still focus on ontology and components, and most companies that truly train, data, and model around the brain are still in the primary market. This is also the reason why he is relatively cautious about the humanoid robot sector in the secondary market.
The aforementioned analyst believes that the current A-share humanoid robot targets are mainly ontology companies. After Yushu Technology goes public, the sector will have a valuation benchmark, and market sentiment will gradually improve. But how far a company on the ontology industry chain can ultimately go depends on its brain. If the brain is not good, the ontology is just hardware, and the brain and data chain are the focus of its next stage of attention.
The breakthrough of brain ability is the prerequisite for everything to be implemented, "Chen Feng said.