Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Generalist AI releases GEN-1.5 robot foundation model that learns new skills from 3-second demos

Robot startup Generalist AI has released GEN-1.5, a foundation model that learns new manipulation skills from just 3-12 second demonstrations, with zero gradient updates or fine-tuning. The model can also chain multiple demos into continuous tasks and transfer virtual demonstrations to the real world, prompting comparisons to a GPT-3 moment for robotics.

Published
看3秒演示就能学会新动作,Generalist AI发布机器人基础模型GEN-1.5
Image source: qbitai.com

Robotics startup Generalist AI has released GEN-1.5, a new robot foundation model built around one-shot learning: a robot that watches a 3-12 second demonstration can immediately perform a brand-new task, with zero gradient updates and no fine-tuning.

The model goes beyond single-task imitation. It can stitch two different teaching demonstrations into one continuous task, filling in transition movements that were never shown — repositioning, adjusting its grasp, switching poses, and even recovering after an error.

Demonstrations do not even have to come from a human. Scripted policies, reinforcement learning agents, and human teleoperation inside a simulator can all be fed into the model's context as physical prompts. After watching a virtual demonstration, GEN-1.5 reproduced the task 1:1 in the real world, despite never having been trained on that task in either setting.

In tests across 10 tasks, GEN-1.5 achieved an average success rate of 59% with a single physical prompt and zero gradient updates; adding five minutes of data and 10 gradient steps per task lifted success to 83%. The team also notes that current test tasks are mostly short-horizon, relatively atomic operations, and skills learned in-context are not yet as stable as those from true fine-tuning.

What excites researchers is that Generalist says it did not design a special architecture, meta-learning loop, or auxiliary objective for one-shot learning — the capability emerged on its own during more than eight months of pre-training on large-scale physical data. During pre-training, the data needed for new tasks shrank from hundreds of fine-tuning steps to dozens, then ten, and finally to just one minute of data and a single gradient step.

That is why the release is being compared to the GPT-3 moment for robotics: GPT-3's in-context learning let models pick up new tasks without retraining, and GEN-1.5 brings the same logic into the physical world, where a demonstrated action replaces a text prompt as the robot's context.

The open questions now are whether zero-training skills stay reliable on longer-horizon, more complex tasks, and whether learning a skill by watching once becomes the standard paradigm for robot foundation models.

Why it matters

If zero-training skill acquisition holds up on longer tasks, it could sharply cut the cost of teaching robots new skills and make one-shot learning a standard feature of robot foundation models.

Generalist AIGEN-1.5Robotics
Back to AI Daily

Nearby Updates

All

08/21, 15:03

MiniMax loses key engineering lead after M3 launch: Skyler Miao departs

MiniMax's Agent Engineering department lead, Ada (Skyler Miao), has left the company, according to Blue Whale News, which cited his Feishu status now showing departure; his next destination is not yet known. Miao described himself on X as MiniMax's Head of Engineering, responsible for M2.x, Agent, Audio, and Hailuo AI, joining in 2023 after stints at Baidu, Beike, and ByteDance.

08/21, 16:03

Ping An reports AI agents drove 57.3B yuan in sales in H1 2026 as daily token usage hits 120B

Ping An has released its first-half 2026 results, reporting that AI agents helped generate more than 57.3 billion yuan in sales. The company's AI systems now consume 120 billion tokens per day, a scale that shows large models deeply embedded in its insurance and financial operations.

08/21, 13:59

SenseTime open-sources 8B multimodal model with native 4K image output

SenseTime has open-sourced an 8B-parameter multimodal model with native 4K image output. The release brings ultra-high-resolution image generation into the open-source community and gives developers a new multimodal foundation model option.

08/21, 16:59

Tencent Cloud and Elastic deepen strategic partnership for AI-era context infrastructure

Tencent Cloud and Elastic have upgraded their strategic partnership to jointly build "context infrastructure" for the AI era. The deepened cooperation targets the growing demand from AI applications for data retrieval and context capabilities.