Economic Observer Follow
2026-10-08 19:12

At the end of August this year, Shandong Jixiang Technology Co., Ltd. (hereinafter referred to as "Jixiang", 6636. HK) submitted its first semi annual report after going public, with a total revenue of 165 million yuan, a year-on-year increase of 181.2%, and adjusted net profit achieving a turnaround. Among them, the revenue from large model solutions has surged more than 45 times year-on-year, becoming the core engine driving performance.
Extreme Vision is known as the "first stock of AI visual big models" and is a representative enterprise of Shandong's artificial intelligence (AI) vertical big models. The 2026 Artificial Intelligence Industry Conference will be held in Jinan, Shandong from October 14th to 16th. Recently, a reporter from the Economic Observer walked into Extreme Perspective and interviewed its co-founder and CEO, Chen Shuo, to learn about the story behind the company's rapid growth.
In 2015, Extreme Vision was founded in Shenzhen by Chen Zhenjie, Chen Shuo, and Luo Yun. This enterprise's breakthrough in the fiercely competitive field of artificial intelligence (AI) vision relies on several key strategic choices.
Chen Shuo introduced that the first key decision is to turn our attention to the real economy fields such as production and manufacturing, target the massive fragmented needs of various industries, and create an AI visual algorithm mall.
With the accumulation of thousands of algorithms and practical experience serving thousands of customers, Jijian has survived the chaos of visual AI and become one of the few AI visual enterprises in China that has achieved profitability.
The second time, I didn't follow the trend of "big model fever" and instead turned around to develop my own interstellar visual language big model, allowing natural language to directly generate recognition models; At the same time, we will create a collaborative application mode and systematic capability for both large and small models, ensuring real-time business response and recognition accuracy with small models, and completing complex understanding of generalized scenarios with large models.
And the second thing cannot be achieved without the scene. Scenarios are the key to realizing commercial value and breaking through technological bottlenecks in large models.
In 2021, Extreme Vision relocated its headquarters from Shenzhen to Qingdao. Regarding the reasons for relocation, Chen Shuo explained that artificial intelligence technology must ultimately be rooted in the context of physical industries in order to realize its value. Shandong has a complete industrial system and concentrated physical enterprises, which can provide a large number of application scenarios and give AI algorithms a realistic application soil. After relocation, the real working conditions from the front line of the industry continue to provide feedback for algorithm refinement and iteration.
On March 30, 2026, Jijian Vision landed on the Hong Kong Stock Exchange, becoming the first Hong Kong Stock Exchange IPO company in Shandong Province in 2026.
In Chen Shuo's view, going public is just a new starting point. He regards the interstellar visual language model as the technological foundation of the company, and in the future, Extreme Vision will anchor two major directions: on the one hand, increase research and development investment to solidify the underlying technology and model platform capabilities; On the other hand, we will explore the second curve and tackle the large-scale application of intelligent agents within enterprises.
Anchoring the visual track
Why is the Extreme Perspective anchored to the field of computer vision? Chen Shuo believes that 80% of human information acquisition comes from the eyes. Looking at the reality of the industry, the visual related infrastructure is already very complete, with surveillance cameras, industrial cameras, drones, and robots covering various industries. This means that Extreme Perspective does not need to build a hardware system from scratch, and can only rely on existing equipment to stack algorithm capabilities to explore the industrial value of AI.
But having complete infrastructure does not mean that demand is easily met. The overall size of the AI vision market is considerable, but the industry demand presents a highly fragmented feature, with numerous segmented and customized business scenarios arising from various industries. It is difficult for a single enterprise's research and development capabilities to cover all market demands.
At that time, most traditional visual AI companies turned their attention to high-value scenarios such as public security and banking, with a focus on areas such as facial recognition and identity verification. The extreme perspective takes a differentiated development path, focusing on the real economy fields such as production and manufacturing, facing massive fragmented industrial scenarios, and attempting to meet 80% of algorithm application needs in the market through a scaled model.
Fragmentation is a challenge for individual companies, but an opportunity for platform models, "said Chen Shuo. In 2017, Extreme Vision launched the first AI visual algorithm mall in China, similar to Apple's App Store. Extreme Vision provides open tools and platforms, gathers global developers, and develops algorithms for thousands of industries. The demand side can acquire AI capabilities just like choosing an app, eliminating the cycle and cost of developing from scratch.
As of now, this mall has listed more than 1500 algorithms, covering over a hundred industry scenarios such as industrial safety production, smart retail, and urban governance. The mall has served over 3000 customers and delivered more than 6000 projects since its establishment, with a product repurchase rate of over 80%.
In this mode, Extreme View does not need to bear all the research and development costs of algorithms alone. The advantages of the light asset model are clearly reflected in financial data: in 2024, the revenue of Extreme Vision was 257 million yuan, a year-on-year increase of 100.8%, the gross profit margin increased to 40.2%, and the adjusted net profit under non international standards turned positive for the first time to 20.5 million yuan.
However, the large number of algorithms produced by the mall are mostly small models trained for a single task, which can only solve specific and isolated business problems. For every new business scenario added, developers have to redo the entire process of data collection, annotation, and model training, which is time-consuming, costly, and difficult to scale up to meet the constantly emerging new demands of customers.
Chen Shuo gave an example: Today, the customer needs to identify safety helmets and reflective clothing, and deploying three small models can meet the requirements; Tomorrow, there will be a demand for identifying black safety helmets in the business, as the model has not undergone corresponding sample training before, and the entire development work will have to start from scratch.
This "one matter, one discussion" development model is becoming increasingly difficult in the context of explosive growth in AI demand. In addition, the demands of industrial customers are also upgrading. From the initial question of 'can you recognize this object' to 'can you understand this scene'.
In order to solve the efficiency bottleneck of traditional algorithms, in 2025, Extreme Vision independently developed a large-scale model of interstellar visual language. The technological breakthrough of this model lies in the fact that over 80% of visual scenes do not require extensive annotation training like traditional solutions. Enterprises can generate customized recognition models with just one click through natural language instructions. This model can adapt to over 100 industry scenarios, greatly reducing the threshold for enterprises to apply large models.
Not being carried away by the 'big model fever'
Compared to pure language models, the development of interstellar visual language models is much more bumpy. Chen Shuo stated that there are numerous cases in the industry where language modeling technology can serve as a reference, but the mature paths that visual modeling can draw on are very limited, and a large number of technical aspects require the company to explore from scratch.
In addition to the challenges at the research and development level, the illusion problem in the implementation of the interstellar visual language big model industry is also quite tricky. Large models often experience inaccurate image understanding and frequent detection errors due to hallucinations in complex real-world scenarios, making it difficult to support core business decisions.
The extreme perspective approaches the problem from two aspects: at the data level, based on over 1 billion real business datasets for annotation training, ensuring high-precision recognition and stable inference of the model in complex scenarios; At the technical level, establish multi-dimensional specialized technical mechanisms such as fine-grained alignment and negative sample sampling to further suppress illusions.
Chen Shuo frankly stated that the problem of illusion is a continuous challenge for all large model enterprises. Large models are essentially probabilistic systems and can never achieve 100% accuracy. Therefore, it is necessary to establish a comprehensive evaluation system to constantly test the boundaries of the model's capabilities "like an exam".
The extreme perspective did not follow in the heat of the big model. The industry generally regards 2023 as the "year of big models". In 2023, Meta launched the SAM model, which is regarded as the "GPT moment" in the field of computer vision, marking the beginning of the era of CV big models. For a time, the competition of large model parameters became the mainstream narrative in the industry.
Chen Shuo believes that in the field of industrial AI, the relationship between large models and traditional small models is not a substitution, but rather a collaborative symbiosis. Based on this judgment, the extreme perspective has chosen a differentiated path of collaboration based on the platform ecosystem undertaking the size model.
Traditional small models are precise and efficient in specific scenarios, but have poor generalization ability. Faced with new scenarios, it is necessary to re collect data, annotate and train, which has long cycles, high costs and low efficiency. Although the general large-scale model can "see", it lacks fine-grained perception and deep understanding of industrial scenarios, and is prone to misjudgment in complex environments such as industry and transportation, making it difficult to support core business decisions.
This collaborative logic has been validated in the cooperation between Extreme Perspective and a large aluminum industry group. Chen Shuo gave an example: The production environment of this enterprise is high temperature, high dust, high corrosion, high noise, and the equipment has a large operating load, resulting in multiple safety hazards such as mechanical injuries. The traditional manual inspection mode is difficult to cover all long tail risks, and also faces problems such as scattered algorithm models, difficulty in adaptation, and difficulty in overall operation and maintenance.
Extreme Perspective takes the Polar Star platform as its management tool, based on the interstellar visual language large model, and has applied 50+models in various scenarios, including 30+large model tasks such as slurry leakage in grinding machines and foreign object detection in pneumatic clutches. The effect shows that the detection rate of artificial defects is only 50% to 60%, while the collaborative mode of large and small models can increase the detection rate to over 95%.
All the algorithm development done by the company is not done in isolation, but comes from every communication with customers, constantly exploring frontline needs. Artificial intelligence needs to be deeply integrated with the real economy in order to unleash maximum value, "said Chen Shuo.
Intelligent agents are a new growth point
Standing at a new starting point after going public, the landing of large-scale enterprise level intelligent agents has become the core direction of the next stage of attack for Extreme Vision.
An intelligent agent is an intelligent system that can autonomously perceive the environment, remember information, make logical decisions, interact with the outside world, and call tools to perform tasks. It is an important form of artificial intelligence products and services, capable of autonomously breaking down complex tasks and achieving specific goals like "digital employees" without continuous human intervention.
Industry data confirms the explosive potential of enterprise level intelligent agents. According to IDC data, the market size of enterprise level agents in China is expected to reach approximately 21.2 billion yuan in 2025 and 44.9 billion yuan in 2026, with a compound annual growth rate of over 110%.
In Chen Shuo's view, many current general intelligent agents can only read documents and complete dialogue and question answering, making it difficult for them to truly intervene in the actual production execution of enterprises, due to the lack of visual perception ability for physical scenes. However, a large amount of business for physical enterprises such as industry, rail transit, and ports occurs on real sites, and pure text models cannot obtain on-site visual information, resulting in general intelligent agents being only able to "answer questions on paper" and unable to truly participate in production execution.
The visual language model is the key to filling this gap. After the intelligent agent connects the business rules and on-site images, it can complete the complete business loop from perceiving the scene, assessing risks, to outputting disposal suggestions. This is precisely the irreplaceable value of visual models. Documents can only tell you 'what should be done', while visuals can tell you 'what happened on site'. Only by combining the two can AI truly work.
Chen Shuo believes that in the future, enterprises will not rely on a single large model to handle all business, but will instead build a set of diverse collaborative model matrices. Different intelligent agents flexibly call corresponding capability bases according to task requirements. Some intelligent agents use visual models to parse on-site images, while others rely on language models to process text business processes. Various models work together to complete complex industrial tasks.
This judgment has been validated through practical application. In 2025, the Polar Perspective Interstellar Visual Language Large Model will serve as the visual capability foundation in the Qingdao Metro Group's large model system, participating in the upgrade planning of over 500 intelligent agents. The related achievements have been put into trial operation on Qingdao Metro Line 6 and other lines.
Taking the pilot of power supply intelligent agents as an example, it can achieve automatic work order distribution, risk warning, fault self diagnosis and prediction. The pilot has shortened the operation time from 12 hours to 3 hours, and improved the disposal efficiency by more than 70%.
In September 2026, Jijian Vision signed strategic agreements with Jihu GitLab and Best Management Consulting to jointly launch the "AI Intelligent Agent Development Alliance". From single point algorithms to intelligent agent ecosystems, Extreme Perspective is completing a strategic leap. The company's positioning has shifted from "AI model capability competition" to "enterprise level AI landing services".
Chen Shuo stated that in the future, in addition to strengthening its capabilities in large-scale modeling, the company hopes to explore how to achieve large-scale deployment of hundreds, thousands, or even tens of thousands of intelligent agents within the enterprise, helping customers achieve organizational change and efficiency revolution through AI. This is expected to become the second curve of the company's future development.

The Central Political and Legal Affairs Commission has released the list of brave warriors for the third quarter of 2026, with Meituan riders Guo Yabin, Yu Rongxian, and Cheng Jianfeng on the list

New "Three Piece Set" in Shopping Mall: Market, IP First Exhibition, First Store

Up to 4.35%! Which bank in mainland China or Hong Kong has a higher fixed deposit interest rate for US dollars?