Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

vivo Unveils BlueLM End-Cloud Model Matrix and Agent-Native Harness Platform at VDC 2026

At the AI session of the 2026 vivo Developer Conference, vivo laid out a 'large models plus Harness' path for building a personal AI assistant, combining an end-cloud BlueLM model matrix with an Agent-native, system-level developer platform. The company says the goal is a phone that understands the present moment, remembers the past, and finishes cross-app tasks under user authorisation and privacy protections.

Published
vivo发布蓝心大模型端云矩阵与Agent原生Harness平台,押注个人专属AI助理
Image source: qbitai.com

At the AI session of the 2026 vivo Developer Conference (VDC), vivo set out a technical path it summarises as "large models plus Harness" — a personalised intelligence foundation intended to rework the operating-system experience. According to the report from QbitAI, the goal is not a model that answers better but a personal assistant that actually gets things done, and a phone that genuinely understands its owner.

The problem vivo framed is a familiar gap: on a phone, stronger AI capability does not automatically mean a better user experience. Model capabilities keep rising and benchmarks keep climbing, yet handset experience has not upgraded by itself. A flight-delay notification can be summarised, but deciding whether to rebook, whether a hotel booking is affected and whether tomorrow's meeting can still be made still falls back on the user.

vivo's answer is a matrix rather than one stronger model. Recognising speech and understanding semantics, tone and intent go to an end-to-end speech model, BlueLM-Realtime; the perception and memory that must stay resident on the device go to the on-device model BlueLM-Nano; high-frequency scenarios such as information lookup, schedule management and cross-app actions go to the cloud-side BlueLM-Flash; and complex reasoning and long-horizon planning are handed to BlueLM-Pro. In specialised domains, the system can also call in external frontier models automatically.

That design moves models from products users must choose into infrastructure the system schedules, so that several models each do their own job under automatic end-cloud coordination. The analogy vivo uses is heterogeneous computing in chips: rather than asking the strongest unit to do everything, let the most suitable unit do the most suitable work. The split is necessary because heavier reasoning costs more money and time, compact on-device models are fast and light but hit their ceiling earlier, and a single real task routinely spans both ends.

The second problem is that people themselves are a constantly changing context. vivo concentrates on perception and memory. For perception, BlueLM-Nano uses a lightweight vision encoder and UI-specific tokens to read screens, images, video and text; for memory, SwiftKV, hybrid sparse attention and PLE layer-wise embeddings extend the on-device context to 32K. All of it sits on the premise of user authorisation and privacy protection, and the wider the sensing and the longer the memory, the heavier the pressure on permission management.

Beyond models, vivo hands the engineering loop to Harness. The "Agent-native BlueLM system-level Harness developer platform" shown at the event has three layers: a perception and memory layer that understands the user and the environment and lets different agents share long-term context; a planning layer that handles routing, task orchestration and reflective optimisation while continuously adjusting execution paths, tokens and cost; and an execution layer that uses more than 6,000 system-level tools plus MCP, A2A, CLI and a unified API to connect models, agents, skills and external services to the phone, to IoT and to the physical world.

Harness is also meant to solve how experience is retained. Built on an observe-reflect-consolidate loop, the system can keep learning from behaviour and feedback under user authorisation and replay task paths while the device is idle, folding the lessons back into system capabilities. Unified intent, a cross-device runtime and context that travels with the user let tasks continue across phone, PC and tablet, with MCP and A2A relaying between agents, models and tools.

In the ecosystem layer vivo positions itself as a bridge rather than a provider of every service. With vivo's assistant Xiaowei (蓝心小V) upgraded to a Pro mode, a developer integrates once and the capability can then be understood, scheduled and reached consistently across entry points. According to the report, vivo and Alipay restructured terminal services into service-direct, service-execution and MCP Skill layers; Meituan connected food, entertainment and lifestyle capabilities to Xiaowei, linking previously scattered service entry points; Amap supports navigation, route planning, local services and POI lookup from a single spoken request; and JD connected cross-app selection, multimodal recognition, product search and payment.

vivo also applied the stack to users who are easily overlooked: since the 2021 "Sheng Sheng You Xi" accessibility programme it has opened 27 accessibility features covering more than 1.1 million users with disabilities, and its multimodal sign-language model, trained on more than 8,000 national-standard words and a million video clips, has moved from single-gesture recognition to continuous sign-language understanding for live translation and learning. The real test for a handset maker is not parameter counts but whether systems engineering turns capability into daily experience — which makes the number of third-party services that actually run on the Harness platform the metric to watch.

Why it matters

The announcement shifts competition from single-model capability toward system-level orchestration and ecosystem on-ramps. If the Harness platform genuinely accumulates developer capability, the moat in handset AI moves from benchmark scores to how much the system can finish on the user's behalf.

vivoAgentOn-device AI
Back to AI Daily

Nearby Updates

All