Realtime AI News
Caffe creator Jia Yangqing launches new AI venture, accelerates GLM-5.2 inference by 534%
Jia Yangqing, creator of the Caffe deep learning framework and former Alibaba VP, has started a new AI optimization venture. His team achieved a 534% inference speedup on Zhipu AI’s GLM-5.2 model, marking a major entry into China’s inference optimization landscape.
Jia Yangqing, the renowned AI infrastructure leader who created the Caffe deep learning framework and later founded Lepton AI (acquired by NVIDIA), has embarked on a new entrepreneurial venture. According to reports, his newly assembled optimization team has achieved a remarkable 534% inference speedup on Zhipu AI’s GLM-5.2 large language model.
Jia’s career trajectory reads like a who’s-who of AI infrastructure. He created Caffe at UC Berkeley, led AI platforms at Facebook, served as Vice President at Alibaba overseeing cloud AI, and founded Lepton AI in 2023 to simplify AI model deployment. After NVIDIA acquired Lepton AI in 2025, Jia joined NVIDIA as VP of System Software. His return to entrepreneurship signals that he sees untapped opportunities in the model optimization layer.
The 534% speedup on GLM-5.2 is particularly significant. GLM-5.2 is Zhipu AI’s latest generation large language model and a key player in China’s domestic LLM landscape. Achieving over five times the inference speed without sacrificing model quality would dramatically reduce deployment costs and improve user experience for enterprises running GLM-based applications.
This development sends several signals about the direction of China’s AI industry. First, inference optimization is emerging as the next major technical battleground. As foundational model capabilities plateau, the ability to run these models faster and cheaper becomes the decisive competitive advantage. Second, a figure of Jia’s stature choosing to start another company in this space validates the thesis that inference optimization represents a massive market opportunity.
Notably, Jia chose GLM-5.2 — a Chinese domestic model — as the demonstration target rather than foreign alternatives. This underscores the maturation of China’s domestic AI ecosystem, where homegrown models and optimization toolchains are forming a complete technical loop. The industry will be watching closely for further details on Jia’s new company, its team composition, funding, and whether it will launch commercial inference optimization products.
Why it matters
Jia Yangqing’s new venture with a 534% GLM-5.2 speedup signals that inference optimization is the next frontier in China’s AI industry, and that domestic model ecosystem is building a complete technical loop.
Nearby Updates
All07/30, 05:07
Thinking Machines co-founder Lilian Weng departs citing health reasons, rejoins OpenAI
Lilian Weng, co-founder of Thinking Machines Lab, stepped down this week citing health impacts from startup pressure, only to rejoin OpenAI where she will lead a top-level research team focused on recursive self-improvement. The move highlights the fierce talent war in AI and the strategic importance of self-improving AI systems.
07/30, 06:35
Tether Data Debuts 460M-Parameter Vision Model, Pushing AI Off the Cloud
Tether Data, the AI arm of the stablecoin issuer Tether, released a 460-million-parameter vision model designed for on-device inference rather than cloud-based processing. The launch marks the company's transition from AI exploration into product delivery.
07/30, 04:52
OpenAI Launches Free AI Access Program for Scientists, Model Weights Remain Off-Limits
OpenAI has announced a new program offering free AI model access to scientific researchers through an application process. However, the company maintains tight control over model weights, continuing its cautious approach to openness and safety.
07/30, 06:46
Microsoft Reports $3.2B Return on Anthropic Investment, OpenAI a 'Mixed Bag'
Microsoft disclosed in its fiscal Q4 2026 earnings that its investment in Anthropic has generated $3.2 billion in returns, while its OpenAI investment yielded mixed results. The financial detail offers a rare glimpse into the performance of the tech giant's dual bets on two competing AI labs.