Realtime AI News
At ECCV 2026, a workshop turned "AI doing business" into an evaluation task
At ECCV 2026 in Malmo, the MARS2 Workshop led by Chinese marketing technology company Tec-Do focused on Agentic Commerce, with a companion challenge drawing 64 teams and more than 1,060 submissions for a $100,000 prize pool. Top scores across the three tracks were 86.67, 62.27 and 63.53, and organizers opened the M-CAR dataset of 3,108 ad videos plus a technical report.
ECCV 2026, the European Conference on Computer Vision, is being held in Malmo, Sweden from September 8 to 12. According to a report from QbitAI, one workshop stood out for its subject matter: MARS2, short for "Multimodal Reasoning and Slow Thinking in the Large Model Era: Towards System 2 and Beyond," which put Agentic Commerce on the agenda.
The report cites ECCV's public information in saying MARS2 was the only Agentic Commerce workshop at this year's conference led by a Chinese technology company - Tec-Do, a cross-border marketing technology firm. The program featured researchers including Oxford professor Yarin Gal, MIT assistant professor Paul Pu Liang, Queen Mary University of London's Shanxin Yuan and Linkoping University's Fahad Shahbaz Khan, discussing how multimodal AI moves from seeing to reasoning.
Alongside the workshop ran the MARS2 2026 Challenge, with a $100,000 prize pool that drew 64 teams and more than 1,060 submissions, according to the report. The task asked whether AI can understand marketing videos, broken into three layers: understand the whole video, locate the key moment, and infer the marketing intent behind it. Each layer mapped to one of three competition tracks.
The three tracks shared a dataset called M-CAR, built from Tec-Do's real commercial scenario. It contains 3,108 advertising videos covering more than 30 languages; 36.5 hours of video in total was split into 18,198 semantic segments, roughly one independent segment every seven seconds. Competitors ran models locally and uploaded results to EvalAI for unified scoring and ranking.
On results, the report gives top scores of 86.67 for the MAC track, 62.27 for VTG and 63.53 for MDC. The Boys won MAC and MDC, while Ya PTers took VTG; the report notes that The Boys, a team with ByteDance ties, finished with two first places and one second place.
The challenge imposed hard constraints: models used in MAC and MDC could not exceed 14B parameters, VTG models could not exceed 8B, and only open-weight models were allowed. The report argues this ruled out brute-forcing scores with huge closed models, pushing teams to refine data handling and reasoning within a limited model size.
The report draws three signals from the results. First, capability boundaries: models handle the rough content of an ad, but performance drops sharply when they must connect multiple segments, trace cause and effect, and make a commercial judgment. Second, an optimization lesson: one VTG team added a cross-modal "audio evidence timeline," improving localization accuracy by 16.7 points, while scaling from 4B to 8B parameters actually cost 0.2 points. Third, an execution pattern: winning approaches used proposer-critic verification, coarse-to-fine two-stage localization and duration-based allocation of visual information - described as evidence-chain engineering.
The broader shift is in the role of such corporate-led workshops. The report argues that leading companies in vertical domains are no longer just sending researchers to submit papers, but are setting the research agenda by posing problems, funding prizes and releasing benchmarks, then feeding winning methods back into industry. The report adds that the event design, evaluation results and winning approaches have been compiled into a technical report, with the benchmark and code to be opened to the research community.
Why it matters
The workshop signals that multimodal AI evaluation is shifting from understanding content to making evidence-based commercial judgments: a 20-plus point gap between top scores across tracks shows key-moment localization and intent reasoning remain weak. For companies with limited compute, cross-modal evidence alignment may pay off more than simply scaling model size.
Nearby Updates
All09/10, 16:17
Amap releases ABot-Earth 0.7, billed as the first 3D-native city world model
Alibaba's Amap released ABot-Earth 0.7 on September 10, describing it as the world's first 3D-native city world model. The model is positioned as an entry point for AI to understand the real world.
09/10, 21:00
d Matrix Adopts NVIDIA NVLink Fusion for Rack Scale XPU Deployment
d Matrix Adopts NVIDIA NVLink Fusion for Rack Scale XPU Deployment. AI inference chipmaker d Matrix today announced it will use NVIDIA NVLink Fusion to connect its next generation Raptor XPUs to NVIDIA’s AI infrastructure platform — joining a growing roster of ecosystem partners. By connecting Raptor to NVIDIA NVLink scale up...
09/10, 15:00
OpenAI and GSA expand AI access and cyber defense for US governments
OpenAI and the U.S. General Services Administration will offer eligible federal, state, local and tribal agencies $0 license fees, 50% off usage and expanded cyber defense support. The arrangement is aimed at lowering the cost and friction of public-sector AI adoption.
09/10, 15:00
Introducing ChatGPT for Financial Services
Introducing ChatGPT for Financial Services. Introducing ChatGPT for Financial Services, combining built in financial data and GPT 6 Astra for research, modeling, and client ready materials.