Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Caffe creator Jia Yangqing launches new AI venture, accelerates GLM-5.2 inference by 534%

Jia Yangqing, creator of the Caffe deep learning framework and former Alibaba VP, has started a new AI optimization venture. His team achieved a 534% inference speedup on Zhipu AI’s GLM-5.2 model, marking a major entry into China’s inference optimization landscape.

Published

Jia Yangqing, the renowned AI infrastructure leader who created the Caffe deep learning framework and later founded Lepton AI (acquired by NVIDIA), has embarked on a new entrepreneurial venture. According to reports, his newly assembled optimization team has achieved a remarkable 534% inference speedup on Zhipu AI’s GLM-5.2 large language model.

Jia’s career trajectory reads like a who’s-who of AI infrastructure. He created Caffe at UC Berkeley, led AI platforms at Facebook, served as Vice President at Alibaba overseeing cloud AI, and founded Lepton AI in 2023 to simplify AI model deployment. After NVIDIA acquired Lepton AI in 2025, Jia joined NVIDIA as VP of System Software. His return to entrepreneurship signals that he sees untapped opportunities in the model optimization layer.

The 534% speedup on GLM-5.2 is particularly significant. GLM-5.2 is Zhipu AI’s latest generation large language model and a key player in China’s domestic LLM landscape. Achieving over five times the inference speed without sacrificing model quality would dramatically reduce deployment costs and improve user experience for enterprises running GLM-based applications.

This development sends several signals about the direction of China’s AI industry. First, inference optimization is emerging as the next major technical battleground. As foundational model capabilities plateau, the ability to run these models faster and cheaper becomes the decisive competitive advantage. Second, a figure of Jia’s stature choosing to start another company in this space validates the thesis that inference optimization represents a massive market opportunity.

Notably, Jia chose GLM-5.2 — a Chinese domestic model — as the demonstration target rather than foreign alternatives. This underscores the maturation of China’s domestic AI ecosystem, where homegrown models and optimization toolchains are forming a complete technical loop. The industry will be watching closely for further details on Jia’s new company, its team composition, funding, and whether it will launch commercial inference optimization products.

Why it matters

Jia Yangqing’s new venture with a 534% GLM-5.2 speedup signals that inference optimization is the next frontier in China’s AI industry, and that domestic model ecosystem is building a complete technical loop.

贾扬清GLM-5.2模型加速推理优化AI创业
Back to realtime news

Nearby Updates

All