Realtime AI News
Om AI Open-Sources On-Device Native VLX-Seek 1.5: Small Parameters, Precise Physical-World Perception
QbitAI reports that Om AI Lianhui has open-sourced VLX-Seek 1.5, an on-device native streaming multimodal model in 3B and 10B versions, designed to perceive the physical world from continuous video streams without cloud dependency. In published benchmarks, the 3B variant beat NVIDIA's same-scale LocateAnything-3B across detection, referring-expression, and drone-view tasks, as the company closed a new funding round of hundreds of millions of yuan.
Chinese tech outlet QbitAI reported on August 10 that Om AI Lianhui has officially open-sourced VLX-Seek 1.5, which the company describes as the world's first on-device native model, offered in 3B and 10B parameter versions. Around the same time, the company closed a new funding round of hundreds of millions of yuan led by Qianhai FOF. The report frames the pairing of open source and capital as an early vote for the on-device native route.
The defining idea behind VLX is on-device native architecture. According to the report, the model is not trained in the cloud and then compressed onto devices; instead, on-device compute, latency, power consumption, and deployment cost are treated as design constraints from the start. The distinction, as the article puts it, is between asking how to make a powerful model run on a device, and asking what a model should look like if it is born on a device.
Technically, VLX-Seek 1.5 uses a streaming multimodal architecture: it takes continuous video input, understands it in real time on the device, and outputs decisions dynamically. That contrasts with the conventional vision-language workflow of snapping a photo, uploading it to the cloud, and waiting for inference results. The report argues this shift from photo recognition to video-stream understanding is what lets a small-parameter model hold its own at the same scale as much larger competitors.
In published evaluations, the report compares VLX-Seek 1.5-3B with NVIDIA's same-scale LocateAnything-3B. On general detection (LVIS Mean) it scored 57.5 versus 50.7; on referring-expression comprehension (RefCOCOg test Mean) it reached 80.2; and on drone-view tests (RefDrone) it posted an F1 of 73.2 and accuracy of 58, up 40% and 62.9% respectively. Its object-hallucination metric (FP/GT) came in at 18, far below LocateAnything-3B's 71.3, indicating far fewer false detections of nonexistent targets.
The article reads these results as evidence that on-device native architecture can deliver efficiency advantages in extreme viewpoints and small-target scenarios. Giants solve problems with scale, it argues, while the on-device native route raises the value of each parameter through an architecture better suited to the physical world. Low hallucination rates matter most for high-reliability settings such as security robots and drone inspection, where false alarms are more dangerous than missed detections.
In the broader landscape, the report maps NVIDIA Cosmos as a cloud world-model route, Google Gemini Robotics as cloud VLA, and Tesla Optimus as a hardware-integrated approach, positioning VLX as the answer to a coordinate none of them addressed: a model designed for on-device environments from birth. The company is also building an embodied-agent platform called OmAgent on the VLX foundation to push AI into the physical world.
The combination of open source and fresh capital is likened to the early cloud-computing transition: capital is pricing a new paradigm, and open source is competing for the right to define the standard for future physical-AI infrastructure. As models leave the cloud and enter real devices, on-device native could become the key answer to scaling physical AI, the article concludes.
What to watch next: whether VLX-Seek 1.5's benchmark results can be independently reproduced, how the 3B and 10B versions perform on real robots, drones, and smart terminals, and whether the open-source release attracts enough developers to build an ecosystem flywheel. The competition ahead may not be about who has the most parameters, but about who defines the infrastructure that lets robots, drones, and smart terminals actually run at scale.
Why it matters
Om AI Lianhui is betting that on-device native architecture, not parameter scale, will define physical AI, backing the bet with an open-source release and fresh funding. If its benchmark results hold up to independent reproduction, edge vision models could scale faster across robots, drones, and smart terminals.
Nearby Updates
All08/10, 08:48
不知不觉间,Deepseek已经变成AI“斩杀线”了? 风闻
不知不觉间,Deepseek已经变成AI“斩杀线”了? 风闻. 不知不觉间,Deepseek已经变成AI“斩杀线”了? 风闻
08/10, 10:04
Tencent's WorkBuddy App Update Lets Users View and Launch AI Tasks on Mobile
Tencent's AI work assistant WorkBuddy has rolled out a mobile app update that lets users view and initiate AI tasks from their phones. The move extends task-oriented AI workflows beyond the desktop, allowing users to track progress and start new tasks on the go.
08/10, 06:41
Orla gives AI agents a dollar-denominated budget with a hard worst-case cap
Singapore-based budgeting app Orla has launched agent payments, letting connected AI agents hold their own stablecoin wallets and pay for services within owner-set limits, where the deposit is the entire possible loss. Invoices priced in USDC can now also accept payments from other agents over the x402 standard, with no setup required for sellers.
08/10, 06:30
Meta launches Muse Code, its first AI coding agent, in beta
Meta has launched Muse Code, its first AI coding agent, in beta, officially entering the AI coding tools market, according to Korean tech outlet DigitalToday. The report says the market once dominated by Anthropic, OpenAI and Cursor is now facing a new wave of competition as big tech players including Google and AWS move in.