Realtime AI News
Local LLM StartLux-27B Passes MCP Benchmark Test: Second Overall, First in Multiple Categories
A CAICT inspection report shows StartLux's local model StartLux-V1.0-27B-Preview ranked second overall in the MCP-specific test of the trusted AI large-model benchmark, beating DeepSeek-V4-Flash-0731 with just a 27B parameter count. It ranked first in location navigation and tied or topped DeepSeek-V4-Pro in browser automation and financial analysis.
An inspection report published by the AI research institute of the China Academy of Information and Communications Technology (CAICT), under the Ministry of Industry and Information Technology, shows that StartLux-V1.0-27B-Preview, a local large language model developed by Shanghai-based StartLux, ranked second overall in the MCP-specific test of the trusted AI large-model benchmark. With a 27B parameter count it finished ahead of third-place DeepSeek-V4-Flash, with capability reaching the band of trillion-parameter models represented by DeepSeek-V4-Pro.
The MCP-specific test covers six task categories — location navigation, web search, browser automation, financial analysis, code repository management, and 3D design — plus a comprehensive evaluation, for seven inspection items in total. It focuses on multi-tool coordination, complex task execution, and interaction with real environments, and its indicators partly reference the open-source MCP-Universe project.
Models tested include DeepSeek-V4-Pro (1.6T), DeepSeek-V4-Flash-0731 (284B), Step-3.7-Flash (198B), StartLux-27B-260715 (27B), Qwen-3.6-27B (27B), and AgentCPM-Explore (4B). StartLux-V1.0-27B-Preview scored 39.25 overall to take second place, beating the 284B DeepSeek-V4-Flash-0731 and the 198B Step-3.7-Flash, and it outperformed the same-size Qwen-3.6-27B by 5.34 percentage points.
In individual tasks the model stood out further: it ranked first in location navigation, and tied with or outright beat the 1.6T-parameter DeepSeek-V4-Pro in browser automation and financial analysis. StartLux says it built the model by post-training Qwen3.6-27B as a base with targeted enhancement, relying on a new multi-dimensional, verifiable, and scalable model optimization technique designed by its own team.
The training process adopted an "AI trains AI" (Auto Research) approach, in which the team autonomously runs training experiments and iterates on strategy through feedback. According to the report, StartLux-V1.0-27B-Preview is the first local agent model in China to be post-trained with this method, which represents a frontier direction in global model development.
The result lands amid a broader push into local models: Google launched the Gemma 4 series in April, with the 31B dense version ranking third on an open-source model leaderboard and running locally on consumer GPUs after quantization; Meta open-sourced the 30B local model Muse Glimmer in early August; and NVIDIA released Nemotron 3.5 Lightning, a 30B MoE model that runs directly on local devices such as RTX PCs.
Industry voices are betting on the category. Chen Danian, founder of Shanda Games and chairman of Yusheng Science, recently predicted that local models will take 80 percent of the large-model market within three years, and Roman Orús, co-founder and chief science officer of quantum and AI software company Multiverse Computing, has said the industry is entering a period in which local LLMs become true competitors to cloud services.
StartLux-V1.0-27B-Preview can already run on consumer PCs, and StartLux plans to launch its first generation of local intelligence solutions within the year while continuing to research new architectures such as diffusion language models. With a 27B local model reaching the capability band of a trillion-parameter flagship, the result validates the post-training optimization route and gives tool-calling benchmarks new weight in evaluating agent performance.
Why it matters
A 27B local model matching trillion-parameter flagships on tool-calling tasks strengthens the case for on-device agents and positions MCP-style evaluations as a key yardstick for agent capability.
Nearby Updates
All09/01, 01:57
CCTV-Affiliated Account Attacks Anthropic, Sets Terms for US-China AI Talks
A CCTV-affiliated account published a lengthy attack on Anthropic on August 31, accusing the company of overstepping user data boundaries and demanding Washington meet two conditions before substantive US-China talks on AI safety can begin. The commentary targets Claude Code over data-transmission allegations and argues US safety rules must first bind American AI companies.
09/01, 01:19
TimesFM 3: A zero shot foundation model for multivariate forecasting
TimesFM 3: A zero shot foundation model for multivariate forecasting. Data Management
09/01, 00:45
Open-source tool weekly-git-report uses MCP and LLMs to turn Git commits into work reports
An open-source report generator called weekly-git-report is gaining attention for integrating the Model Context Protocol and using an LLM to automatically convert Git commits into work reports. It is a concrete example of AI agents moving into routine developer workflows.
09/01, 00:00
AI video search startup Clipto raises $15M at a $250M valuation
Three-year-old AI media search startup Clipto has raised $15 million in an all-equity round at a $250 million post-money valuation. The company says it hit $15 million in annual recurring revenue and remains profitable on a net-income basis.