Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

AgentGarten links code-built worlds with real-time neural rendering for self-evolving agents

The MirroS team has released AgentGarten, an open-source framework that pairs executable code environments with a real-time neural renderer so agents can act, observe and improve above 30 fps. In a hide-and-seek test, agents learned to build shelters and climb ramps within a few rounds, using notebooks they wrote themselves.

Published
代码造世界、扩散绘现实:AgentGarten 让智能体在实时试炼场里边玩边进化
Image source: qbitai.com

Push a board into place and a corridor closes; drag a ramp over and a high wall becomes climbable. Every action changes the world, and the changed world becomes the starting point for the agent's next decision. AgentGarten, newly released by the MirroS team, is an attempt to build AI a training ground it can explore over and over.

The idea is simple to state: physics goes to code, light and shadow go to a neural model. After the agent issues an action, the code environment immediately computes collisions and state updates and exports a lightweight geometric sketch — spatial depth or surface normals — from the agent's first-person camera. A neural renderer then reads that sketch, combines it with visual memory, and generates a photorealistic next frame for the agent.

The team argues both mainstream approaches fall short. Traditional game engines offer precise, transparent states, but their visuals are often coarse white-box geometry and repeating textures, and scaling them to hundreds of realistic scenes demands unaffordable art and engineering. Video-generation world models produce striking imagery, yet their physics stay implicit inside network memory, with no queryable or intervenable state. AgentGarten splits the work so it keeps physical determinism while gaining realistic rendering.

That division changes two things. Physical rules stay deterministic and controllable — scene layout, contact decisions and win conditions are all adjudicated by code, with no "hallucinated physics." And building new worlds becomes far faster: any minimal code scene that can output depth and normal outlines can reuse the same neural renderer. Wingsuit flight, robotic-arm manipulation, kitchen cooking and multi-car racing share one visual interface, shifting the cost of scaling virtual environments from labor-intensive art to programmatic code.

With the world in place, the team tested whether agents can improve on their own using the classic hide-and-seek setup. Hiders first block entrances, then seekers enter. Both sides write Python code to control movement and grabbing, deciding purely from the rendered frames in front of them and knowing nothing about spatial coordinates or the opponent's position. Starting from blank notebooks, each side maintains a strategy library of single skills, plays ten games per round, reflects, and draws part of the library at random next round.

Progress came quickly. Hiders learned to move boards into shelters by round four, and seekers learned to use ramps to climb walls by round ten; after a failed jump, a seeker would pull the ramp closer and try again. For comparison, in OpenAI's well-known 2019 study, agents using from-scratch reinforcement learning took roughly 25 million games to discover shelter building and about 100 million games to learn ramp climbing.

What drives the evolution is a set of experiment notebooks the agents write themselves. At the end of each round they review attempts, record findings and flag open questions, carefully separating what they actually saw from what they merely guessed, and leave anti-pitfall notes for successors. Lessons that survive scrutiny carry forward, while failed ones push them toward new branches. The team split practice into four progressive stages and ran four rounds each across four very different worlds to check the loop generalizes.

To support acting while watching, the neural renderer still had to clear three hurdles: keeping up with sudden actions, staying stable over long interactions and running fast enough. From a foundational omni-modal model, the team proposed an Adversarial Forcing training method, converted whole-clip offline generation into streaming, chunk-by-chunk rendering, designed an exact-replay mechanism to curb appearance drift over long sequences, and added a real-video discriminator as an adversarial signal to prevent visual degradation. On the engineering side they hand-wrote Triton fused operators, used CUDA Graph to cut scheduling overhead, and added a lightweight decoder with under 10 ms latency, sustaining closed-loop interaction at 480p and above 30 fps.

The team calls the concept Physical RSI — recursive self-improvement in the physical world. Executable code generates worlds endlessly, the learned renderer gives those worlds perceptible physical feedback, and the accumulated notebooks make each attempt a step for the next. The code is open-sourced on GitHub, with a technical report and project page. In the team's view, scaling for language models is in full swing, while scaling for intelligence in the physical world has only just begun.

Why it matters

By separating physical determinism from neural rendering, AgentGarten sidesteps the inability of video world models to expose an intervenable state, offering embodied AI a lower-cost path to scale simulation. Whether it can move from hide-and-seek to real robot training is the next thing to watch.

AgentOpen SourceEmbodied AI
Back to realtime news

Nearby Updates

All

10/09, 09:53

After Korean bank hack, ARTEX developer takes the AI agent closed-source

The developer behind ARTEX, an open-source AI penetration-testing tool from China, has converted the project to closed source after CrowdStrike linked it to attacks on South Korean financial firms. The developer denies involvement and says the tool was meant for authorized security testing.

10/09, 10:37

openJiuwen Open-Sources Enterprise-Grade AgentOS, Betting on Self-Evolving Multi-Agent Teams

On October 9, openJiuwen released and open-sourced an enterprise-grade AgentOS, aiming to move AI agents from isolated demos to large-scale enterprise deployment. The company highlights multi-agent collaboration and a self-evolving capability designed for complex business tasks.

10/09, 08:44

Anthropic releases a free AI security scanning service for open-source projects

Anthropic has released a free AI-powered security scanning service aimed at open-source projects, giving maintainers a no-cost way to check their code for vulnerabilities. The move points frontier AI labs toward the open-source software supply chain, a longstanding weak spot in how modern software is built and shipped.

10/09, 08:35

Terence Tao leads mathematicians' association in a joint boycott of OpenAI

Fields Medal winner Terence Tao is leading the Association for Human Mathematicians (AHM) in a joint statement urging mathematicians worldwide to stop collaborating with OpenAI and boycott the company. The statement targets a batch of machine-generated manuscripts OpenAI published, calling it not research but a display of compute power.