Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

At ECCV 2026, a workshop turned "AI doing business" into an evaluation task

At ECCV 2026 in Malmo, the MARS2 Workshop led by Chinese marketing technology company Tec-Do focused on Agentic Commerce, with a companion challenge drawing 64 teams and more than 1,060 submissions for a $100,000 prize pool. Top scores across the three tracks were 86.67, 62.27 and 63.53, and organizers opened the M-CAR dataset of 3,108 ad videos plus a technical report.

Published
ECCV 2026上,MARS2 Workshop把“AI做生意”变成了一道评测题
Image source: qbitai.com

ECCV 2026, the European Conference on Computer Vision, is being held in Malmo, Sweden from September 8 to 12. According to a report from QbitAI, one workshop stood out for its subject matter: MARS2, short for "Multimodal Reasoning and Slow Thinking in the Large Model Era: Towards System 2 and Beyond," which put Agentic Commerce on the agenda.

The report cites ECCV's public information in saying MARS2 was the only Agentic Commerce workshop at this year's conference led by a Chinese technology company - Tec-Do, a cross-border marketing technology firm. The program featured researchers including Oxford professor Yarin Gal, MIT assistant professor Paul Pu Liang, Queen Mary University of London's Shanxin Yuan and Linkoping University's Fahad Shahbaz Khan, discussing how multimodal AI moves from seeing to reasoning.

Alongside the workshop ran the MARS2 2026 Challenge, with a $100,000 prize pool that drew 64 teams and more than 1,060 submissions, according to the report. The task asked whether AI can understand marketing videos, broken into three layers: understand the whole video, locate the key moment, and infer the marketing intent behind it. Each layer mapped to one of three competition tracks.

The three tracks shared a dataset called M-CAR, built from Tec-Do's real commercial scenario. It contains 3,108 advertising videos covering more than 30 languages; 36.5 hours of video in total was split into 18,198 semantic segments, roughly one independent segment every seven seconds. Competitors ran models locally and uploaded results to EvalAI for unified scoring and ranking.

On results, the report gives top scores of 86.67 for the MAC track, 62.27 for VTG and 63.53 for MDC. The Boys won MAC and MDC, while Ya PTers took VTG; the report notes that The Boys, a team with ByteDance ties, finished with two first places and one second place.

The challenge imposed hard constraints: models used in MAC and MDC could not exceed 14B parameters, VTG models could not exceed 8B, and only open-weight models were allowed. The report argues this ruled out brute-forcing scores with huge closed models, pushing teams to refine data handling and reasoning within a limited model size.

The report draws three signals from the results. First, capability boundaries: models handle the rough content of an ad, but performance drops sharply when they must connect multiple segments, trace cause and effect, and make a commercial judgment. Second, an optimization lesson: one VTG team added a cross-modal "audio evidence timeline," improving localization accuracy by 16.7 points, while scaling from 4B to 8B parameters actually cost 0.2 points. Third, an execution pattern: winning approaches used proposer-critic verification, coarse-to-fine two-stage localization and duration-based allocation of visual information - described as evidence-chain engineering.

The broader shift is in the role of such corporate-led workshops. The report argues that leading companies in vertical domains are no longer just sending researchers to submit papers, but are setting the research agenda by posing problems, funding prizes and releasing benchmarks, then feeding winning methods back into industry. The report adds that the event design, evaluation results and winning approaches have been compiled into a technical report, with the benchmark and code to be opened to the research community.

Why it matters

The workshop signals that multimodal AI evaluation is shifting from understanding content to making evidence-based commercial judgments: a 20-plus point gap between top scores across tracks shows key-moment localization and intent reasoning remain weak. For companies with limited compute, cross-modal evidence alignment may pay off more than simply scaling model size.

MultimodalAgentECCV
Back to realtime news

Nearby Updates

All