Realtime AI News
Generalist AI releases GEN-1.5 robot foundation model that learns new skills from 3-second demos
Robot startup Generalist AI has released GEN-1.5, a foundation model that learns new manipulation skills from just 3-12 second demonstrations, with zero gradient updates or fine-tuning. The model can also chain multiple demos into continuous tasks and transfer virtual demonstrations to the real world, prompting comparisons to a GPT-3 moment for robotics.
Robotics startup Generalist AI has released GEN-1.5, a new robot foundation model built around one-shot learning: a robot that watches a 3-12 second demonstration can immediately perform a brand-new task, with zero gradient updates and no fine-tuning.
The model goes beyond single-task imitation. It can stitch two different teaching demonstrations into one continuous task, filling in transition movements that were never shown — repositioning, adjusting its grasp, switching poses, and even recovering after an error.
Demonstrations do not even have to come from a human. Scripted policies, reinforcement learning agents, and human teleoperation inside a simulator can all be fed into the model's context as physical prompts. After watching a virtual demonstration, GEN-1.5 reproduced the task 1:1 in the real world, despite never having been trained on that task in either setting.
In tests across 10 tasks, GEN-1.5 achieved an average success rate of 59% with a single physical prompt and zero gradient updates; adding five minutes of data and 10 gradient steps per task lifted success to 83%. The team also notes that current test tasks are mostly short-horizon, relatively atomic operations, and skills learned in-context are not yet as stable as those from true fine-tuning.
What excites researchers is that Generalist says it did not design a special architecture, meta-learning loop, or auxiliary objective for one-shot learning — the capability emerged on its own during more than eight months of pre-training on large-scale physical data. During pre-training, the data needed for new tasks shrank from hundreds of fine-tuning steps to dozens, then ten, and finally to just one minute of data and a single gradient step.
That is why the release is being compared to the GPT-3 moment for robotics: GPT-3's in-context learning let models pick up new tasks without retraining, and GEN-1.5 brings the same logic into the physical world, where a demonstrated action replaces a text prompt as the robot's context.
The open questions now are whether zero-training skills stay reliable on longer-horizon, more complex tasks, and whether learning a skill by watching once becomes the standard paradigm for robot foundation models.
Why it matters
If zero-training skill acquisition holds up on longer tasks, it could sharply cut the cost of teaching robots new skills and make one-shot learning a standard feature of robot foundation models.
Nearby Updates
All08/21, 15:03
MiniMax loses key engineering lead after M3 launch: Skyler Miao departs
MiniMax's Agent Engineering department lead, Ada (Skyler Miao), has left the company, according to Blue Whale News, which cited his Feishu status now showing departure; his next destination is not yet known. Miao described himself on X as MiniMax's Head of Engineering, responsible for M2.x, Agent, Audio, and Hailuo AI, joining in 2023 after stints at Baidu, Beike, and ByteDance.
08/21, 16:03
Ping An reports AI agents drove 57.3B yuan in sales in H1 2026 as daily token usage hits 120B
Ping An has released its first-half 2026 results, reporting that AI agents helped generate more than 57.3 billion yuan in sales. The company's AI systems now consume 120 billion tokens per day, a scale that shows large models deeply embedded in its insurance and financial operations.
08/21, 13:59
SenseTime open-sources 8B multimodal model with native 4K image output
SenseTime has open-sourced an 8B-parameter multimodal model with native 4K image output. The release brings ultra-high-resolution image generation into the open-source community and gives developers a new multimodal foundation model option.
08/21, 17:01
Anthropic adds AI watermark to Claude, sparking strong user backlash
Anthropic has added an AI watermark to Claude to flag AI-generated content, and the move has sparked strong backlash from users. The controversy highlights the tension between content-provenance requirements and the everyday product experience that users expect.