Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Lenovo's TianxiCode Tops SWE-bench-Live With 71% Resolution Rate

Lenovo's Tianxi AI unit says its self-developed code agent framework, TianxiCode, has taken the global number-one spot on SWE-bench-Live with a 71% problem-resolution rate. The result puts an enterprise-built coding agent on one of the most closely watched public benchmarks for autonomous software work.

Published

Lenovo's Tianxi AI unit has disclosed that TianxiCode, a code-focused agent framework it developed in-house, has claimed the global number-one position on SWE-bench-Live with a 71% problem-resolution rate, according to a report from QbitAI. For an enterprise-built coding agent, the result is a direct signal: this is not a concept demo but a claim measured against a public, continuously refreshed benchmark for real software engineering tasks.

SWE-bench has become one of the most-watched ways to measure what coding agents can actually do. It hands agents real repository issues and scores them on whether they can locate, modify and verify code on their own — in other words, whether they truly fix the problem. SWE-bench-Live emphasizes fresh, dynamic tasks, so a model cannot simply rely on problems it saw during training; it has to keep performing on newly added ones.

TianxiCode is described as Lenovo Tianxi AI's self-developed professional code agent framework, positioned for code work rather than general chat. The 71% resolution rate is the key disclosure here: it determines the framework's ranking and gives outsiders the most direct measure of how finished the engineering is.

The broader context is that coding agents have been among the fastest-moving agent categories of the past year. Unlike conversational skills, code tasks have a built-in correctness test — whether tests pass and defects get fixed — so it is hard to fake with fluent language. That is why a high placing on a benchmark like SWE-bench is often read as evidence that an agent can genuinely carry out work rather than merely describe it.

Lenovo's timing also suggests its enterprise AI strategy is extending from platform and model layers into concrete, high-value agent use cases. For a company with a large hardware, device and enterprise-services footprint, a demonstrable coding agent can create a new entry point into customers' development and operations workflows.

One caveat is that a single benchmark ranking does not equal broad real-world capability. SWE-bench-Live focuses on software engineering tasks, while enterprise deployment also involves private codebases, access control, compliance and cost. Whether TianxiCode can turn a leaderboard result into a product that scales remains an engineering and ecosystem question.

Two things are worth watching next: whether Lenovo publishes TianxiCode's technical details, availability and deployment model, and whether the result pushes more homegrown agents to compete on shared public benchmarks — moving the field from launch-day claims toward reproducible, verifiable comparison.

Why it matters

Topping SWE-bench-Live gives Lenovo a public benchmark credential in the enterprise coding-agent race, useful for pitching developer agents to business customers. The result becomes a durable advantage only if Lenovo follows through with disclosed technical details and a deployable product.

LenovoCoding AgentSWE-bench
Back to realtime news

Nearby Updates

All

10/09, 16:47

Aether AI's CRIS-0 brings causal intelligence to real robots with 0.2-second safety stops

Aether AI, the causal-AI company founded by UCSD professor Biwei Huang, has unveiled a demo of its CRIS-0 robotics system, which re-plans around disturbances such as a shifted coffee machine in about two seconds on average and recovered in 9 of 10 trials. The system models tasks as evolving causal states and can halt a closing microwave door in 0.2 seconds when a human hand appears, targeting the error accumulation and brittleness of end-to-end models.

10/09, 15:52

ICANN publishes application list, with OpenAI seeking .gpt, .chatgpt and .agi

ICANN has published the list of applications for new generic top-level domains, and OpenAI is among the applicants. OpenAI is reported to have applied for the strings .gpt, .chatgpt and .agi, signalling an effort to turn its brand and core concepts into internet addresses.

10/09, 18:00

Google locks Gemini Pro behind its most expensive subscriptions

Google is restricting which Gemini models each plan can use, limiting free users to the weaker Gemini Flash Lite and pulling Gemini Pro from all but its most expensive subscriptions. The change, starting October 9, is aimed at converting free users into paying customers.

10/09, 15:00

Sophos cuts threat investigation time by 96% with OpenAI Daybreak

OpenAI has detailed how the security vendor Sophos uses its Daybreak offering to cut cyber-threat investigation time by 96% and automate 52% of MDR cases. The case highlights agentic AI moving into security operations while keeping humans in the loop.