MirroS
Research·Oct 8, 2026

AgentGarten:
Code Worlds for
Evolving Agents

Executable worlds, open-ended interaction, and a real-time visual interface. Building the environment in which an agent can turn action into experience.

A World to Act In

An agent needs an environment it can interact with. It must be able to choose an action, observe its consequences, and decide what to do next. Turning a corner changes what is visible; moving an object changes what can happen afterward. These exchanges become trajectories of experience. To support exploration and learning, the environment must sustain them across navigation, manipulation, tool use, and repeated attempts at a task.

Code-based environments give us direct control over this interaction. We can specify objects, implement rules, and keep track of changes to the world[1]. Their visual quality, however, depends on the assets and rendering pipeline we build. A procedurally assembled scene may have useful interaction logic while producing simplified geometry, repetitive materials, or unconvincing lighting. Richer visuals are possible, but creating them across many environments requires substantial asset and rendering work.

World models built around video generation[2][3] offer another route to rich visual observations, with camera and action inputs controlling generation. In these systems, information about the world is largely carried through generated history and internal memory. This can support visual continuity, but does not necessarily provide an explicit world state that can be queried and edited directly. Inspecting an object’s attributes, changing a relationship between objects, or specifying how an action should affect future interactions is therefore difficult. Movement and camera controls can support navigation, while richer interactions—such as manipulating objects, using tools, or enforcing task-specific rules—require more direct access to the underlying state and its transitions. Rich visual generation alone does not provide this programmable interface.

We therefore pair a code-based environment with a learned real-time neural renderer: code determines how the world changes, and the renderer learns how those changes should look. Unlike approaches that generate a full video window from conditions supplied in advance[4], our renderer accepts conditions in short chunks and generates observations incrementally. An agent can take a brief action, inspect its consequences, and decide what to do next before the next chunk is generated. Future conditions therefore depend on the agent’s response to what it actually observes. A general interface lets navigation, object manipulation, and multi-agent environments share the same renderer. Each environment retains direct control over its objects, interaction rules, and task goals through code.

We use this loop to study agents that improve from experience. In hide-and-seek, the game of OpenAI’s 2019 study of emergent tool use[5], agents that act only on rendered first-person frames learn to build shelters and use ramps within a few rounds of play and review. The same procedure carries over to other worlds with very different tasks.

Code worldAgentactionsEngineworld = Bridge(road, cars)def step(actions): for car, a in actions: car.drive(a) world.block_collisions() return world.stateConditionsReference + textNeuralrendererVisual historyObservationappendobservationsAgentactionsEngineworld = Bridge(road, cars)def step(actions): for car, a in actions: car.drive(a) world.block_collisions() return world.stateConditionsNeural rendererObservationobservations
Figure 1. Interaction and rendering in a code world. The agent acts on the code world, which updates its state and exports structured conditions. The neural renderer turns these conditions into the agent’s next visual observation, closing the interaction loop.

A Scalable, Versatile World Model

Scaling the environment library starts with code: layouts, objects, dynamics, interaction rules, and task goals can be generated and modified systematically. These environments share a learned renderer, reducing the need to build detailed visual assets and tune a rendering pipeline for each one. This makes it easier to create a larger and more varied set of environments for agent learning.

ConditionRendered
Figure 2. Neural rendering across code worlds. The examples cover navigation, object manipulation, scene re-shooting, and multi-agent interaction.

Agents That Improve from Experience

A world that answers every action is also a place to practice. We begin with hide-and-seek, a game with a history, then run the same round-by-round procedure in four more worlds.

Hide-and-Seek, Revisited

In OpenAI’s 2019 hide-and-seek study[5], agents trained with self-play and reinforcement learning developed strategies such as building shelters, using ramps, and defending against those tools. The paper reports shelter construction after roughly 25 million episodes, followed by seeker ramp use after another 75 million. The environment made a sequence of increasingly sophisticated strategies possible through repeated competition.

We revisit this setting with a pretrained agent[6] that can reason about what happened and revise how it acts. Within several rounds of interaction and review, the agents had already accumulated and refined reusable lessons about object control, shelter access, and ramp traversal.

The game has two opposing roles: hider and seeker. After every round, each role reviews its games and writes down what it learned. Each keeps its own playbook, organized as a library of short skill files with one lesson per file.

Our framework uses a sequential one-hider, one-seeker variant of the physics environment. The hider prepares the scene, then hands control to the seeker. Agents receive first-person visual observations generated by the neural renderer from the environment’s structured conditions. They submit short Python action programs for movement and object interaction, then observe the resulting changes through the renderer. They do not receive object coordinates, hidden world state, or the opponent’s private observations.

The experiment begins with empty playbooks, and each round consists of five games played across sampled layouts and random seeds. After each round, each role reviews its own action programs and permitted visual evidence, identifies failures, and adds skill files to its playbook: lessons or code that may help in later games. In the next round, the agents receive the accumulated skills, interpret the current scene, and determine which prior experience is relevant. Both successful and unsuccessful attempts contribute to subsequent revisions.

Aspect

Self-play RL [5]

Pretrained agents in a world modelOurs

Agent
Policy network
Pretrained agent
Sees
Object state
Agentspositionvelocity
Boxespositionvelocitysize
Rampspositionvelocity
Lidarrange readings
First-person frames
Hider’s view
Seeker’s view
Acts
One action per step
Move x
Move y
Turn
Short Python programs
def policy():
    step(turn=-1, steps=2)
    pull(1)
    goto(180, 2)
Keeps experience in
Policy weights
playbooks
Figure 3. Hide-and-seek in two settings. Self-play RL trains a policy network on object state and rewards. Our pretrained agents see neural-rendered first-person frames, act through short Python programs, and keep what they learn in a playbook of skill files. Click a skill file to read notes from the first-round review.

The clips below are three of several emerged behaviors we observed from this experiment. A hider builds its cover by moving a panel; a seeker carries a ramp to a wall and climbs over it; and, when a climb falls short, a seeker moves the ramp closer and tries again.

Build the cover

The hider carries a long panel to the corner and turns it, changing the barrier layout before the seeker's turn.

Carry a ramp to the wall

The seeker carries a ramp to the inner wall, climbs over it, and reaches the hider on the other side.

Move the ramp and try again

The first climb falls short. The seeker moves the ramp closer, climbs again, and reaches the hider.

Figure 4. Strategies in three recorded games. Each clip shows one role’s phase of the game.

These are the behaviors for which the 2019 study is best known, providing a useful point of comparison (Figure 5). In that study, shelter construction emerged after roughly 25 million training episodes, while ramp use appeared after roughly 100 million. In our setting, the hiders used a panel to build a shelter by round 4, and the seekers used a ramp to cross walls by round 10.

When do strategies emerge?

StrategyOpenAI · self-play RL [5]training episodes Visual agentsrounds played
Build shelters≈25M4
Use ramps to enter shelters≈100M10
Figure 5. Rounds played before each strategy appears. Self-play RL trains policies from scratch on object state over millions of episodes. Our pretrained agents act on rendered first-person frames and revise written playbooks between rounds.

Beyond Hide-and-Seek

The hide-and-seek agents improved by playing, reviewing their games, and writing down what they learned. This loop carries over to other tasks unchanged (Figure 6). A round begins with a task file that states the goal, the available actions, and the limits, but gives no solution. One or more agents then play in parallel, each within a fixed budget of steps or simulated time. They see the world only through camera frames: no coordinates, no map, and no score until the episode ends. Afterwards each agent writes a playbook recording what it tried, what it observed, what it is still unsure of, and what to test next. Playbooks are frozen and archived, and the next round’s agents start from the task file and the playbooks of earlier rounds.

Agent Loop

  1. 1Read

  2. 2Play

    Agent 1
    Agent 2
    Agent 3
    Agent 4
  3. 3Write

  4. 4Archive

Next round: new agents start from the task and earlier playbook
Figure 6. One round of practice. Agents read the task, play in parallel within a fixed budget, and write a playbook; the next round starts from the task and the archived playbooks. The files shown are an example from the one-lane bridge world; click one to read it.

A code world is a program, so the procedure can be pointed at very different tasks. We ran it for four rounds in each of four more worlds: a companion dog, a one-lane bridge, sheep herding, and a quarry loader (Figure 7). Each world has its own actions, time limit, and score. This is what the agents are asked to do in each:

  • Companion dogAn agent has a 60-second session with a dog and must keep it comfortable and willingly engaged. It can offer a hand to sniff, stroke the dog’s chest or head, play with a ball, or draw back. A dog at ease takes small steps toward the agent and stays close; an uneasy one turns aside and steps away. From the camera view alone, the agent has to learn which combinations this dog responds to.
  • One-lane bridgeTwo cars start on opposite banks of a bridge that fits only one. Each is driven by its own agent from the windshield view, and both must reach the other side as quickly as possible, within 120 simulated seconds. Pullouts on the banks let one car wait while the other passes.
  • HerdingTwo dogs, each seeing only from its own eye height, have 180 seconds to guide four sheep into a fenced pen and keep them inside for five seconds. Sheep move away from a nearby dog and scatter when pressed too closely.
  • Quarry loaderA wheel loader has 360 seconds for three jobs in order: push two rocks onto staging pads, move one of them around a wall into a bunker, and park in its bay.

Companion dog

Engagement score

13Round 114Round 219Round 319Round 4

One-lane bridge

Seconds until both cars arrive · lower is better

71 sRound 168 sRound 245 sRound 341 sRound 4

Herding

Score · out of 100

60Round 190.1Round 287.6Round 388.3Round 4

Quarry loader

Score · out of 100

30Round 10Round 230Round 390.9Round 4
Figure 7. Four more worlds, four rounds each. Each card shows the first and the fourth round from the agent’s camera, the world’s own measure by round, and the task file and playbooks from the run.

Building a Real-Time Visual Interface

The neural renderer turns the structured conditions exported by the code world into what the agent sees. To support interaction, it must do three things: follow each new condition, stay consistent over long rollouts, and generate observations fast enough for the agent to respond. We start from a pretrained omni model[7] and adapt it in two steps: geometry conditioning and Adversarial Forcing. At inference, hand-written fused kernels, CUDA graph capture let the renderer run in real time.

Following the Code World: Geometry Conditioning

The code world passes its state to the renderer as geometry: depth or surface normals, seen from the agent’s camera. We chose geometry because it is easy to obtain for both training and inference. For training, simulators export depth and normals directly, and for real videos we estimate them with depth estimators[8] and normal estimators[9]. At inference, any code world with a 3D scene can render them cheaply, without detailed assets or materials.

Adversarial Forcing

Adversarial Forcing turns the geometry-conditioned model into a few-step renderer that stays consistent over long rollouts. It has three parts: teacher-forcing training of a block-causal model, exact replay so that gradients reach the history the model wrote, and distribution matching combined with a real-data adversarial signal regularized by exact R1/R2.

Generating Block by Block

To respond to actions frequently during a rollout, we adapt the bidirectional pretrained model to generate short blocks with block-causal attention: each frame can attend within its block and to preceding blocks. We first train it with teacher forcing, conditioning each block on ground-truth history.

A model trained this way has only seen clean history. At inference, it conditions on its own outputs, and small errors accumulate over long rollouts. Following Self Forcing[10], we train the model on rollouts conditioned on its own outputs and match their distribution to the teacher’s using Distribution Matching Distillation (DMD)[11], which also reduces the number of sampling steps.

Self-forcing DMD on its own leaves two gaps that matter over long rollouts. Gradients do not reach the computation that stored the generated history, and nothing in the objective compares outputs with real videos. We close the first with exact replay and the second with a real-data adversarial signal.

Exact Replay for History Gradients

Standard Self Forcing generates the rollout in a single pass and detaches the generated history to reduce peak memory, but this blocks gradients through the clean-history prefill computation and may contribute to drift. We use two-pass training to make this computation differentiable. The first pass rolls out the video without gradients; the second replays it with gradients through the clean-history computation, using exactly the same attention kernel. Self Gradient Forcing (SGF)[12] makes clean-history prefilling differentiable with a similar two-pass process, but replays the second pass with a FlexAttention[13] kernel, which causes a 1.41% relative L2 error between rollout and replay, as reported in its paper; in our own matched test, an SGF-style full-sequence replay produced 3.99% error. Our replay recomputes the history keys and values within the graph, but runs attention block by block with the same scaled dot-product attention (SDPA) kernel, key/value order and call shapes as the first-pass rollout (Figure 8). In real training configurations, the replay is bitwise identical to the rollout.

Pass 1keys / valuesqueries12341234Pass 2 · Ours12341234Pass 2 · SGF-style12341234read from the detached cachecomputed, no gradientscomputed with gradientsone attention callPass 1keys / valuesqueries12341234Pass 2 · Ours12341234Pass 2 · SGF-style12341234read from the detached cachecomputed, no gradientscomputed with gradientsone attention call
Figure 8. Exact replay, drawn as block-level attention masks. Row i holds block i’s queries, and the columns are the key/value blocks they read. The rollout runs one attention call per block and reads earlier blocks from a detached KV cache. Our replay runs the same calls but computes the history keys and values with gradients, so it matches the rollout bitwise. An SGF-style replay computes the same mask in one call; the different execution order gave a 3.99% relative L2 error in our matched test.

A Real-Data Adversarial Signal

Most self-forcing DMD methods[14, 15, 16] train the student with score distillation alone, using teacher and fake-score estimates evaluated on the student’s own samples. This objective does not directly compare outputs with real videos. In our geometry-conditioned setting, models trained with score distillation alone exhibited collapse during long rollouts and major scene changes: generated frames developed checkerboard-like textures and drifted out of alignment with the conditioning video. Following DMD2[17], we add a GAN loss on real videos. Conditioned on geometry and history, the discriminator penalizes both implausible appearance and deviations from the condition. This additional signal helped preserve texture and condition following across long rollouts.

Exact R1/R2 without Double Backward

We use R3GAN’s relativistic objective with R1/R2 gradient penalties[18]. These penalties normally require differentiating through a backward pass (double backward), which the fused attention kernels we use, such as FlashAttention[19] do not support. APT[20] works around this with a random-perturbation approximation of R1. We compute the full input-gradient penalty and its exact head-parameter gradient instead. The discriminator’s backbone is frozen, so a vector–Jacobian product (VJP) and a Jacobian–vector product (JVP) through it supply the input gradient and its direction in feature space. The penalty then joins the relativistic loss in the discriminator’s ordinary update: one head forward and one backward pass (Figure 9).

xx
FrozenBackbone
BB
zz
TrainableHead
hϕh_\phi
DϕD_\phi
12
vv
3
Lrel+γ2R~\mathcal{L}_{\mathrm{rel}}+\tfrac{\gamma}{2}\tilde R
xx
FrozenBackbone
BB
zz
TrainableHead
hϕh_\phi
DϕD_\phi
12
vv
3
Lrel+γ2R~\mathcal{L}_{\mathrm{rel}}+\tfrac{\gamma}{2}\tilde R
  1. 1Input gradientg=∇xDϕ(x)g=\nabla_x D_\phi(x)
  2. 2Feature directionv=JB(x) gv=J_B(x)\,g
  3. 3Discriminator update∇ϕ[Lrel+γ2R~]\nabla_\phi\big[\mathcal{L}_{\mathrm{rel}}+\tfrac{\gamma}{2}\tilde R\big]
Figure 9. Exact R1/R2 with a frozen backbone. (1) A backward pass gives the input gradient g, whose squared norm is the penalty. (2) A JVP carries g through the frozen backbone to the feature direction v. (3) The head runs once on the detached features and v, producing the relativistic loss and the penalty term together, and one backward pass updates it. No step differentiates through a backward pass.
Exact R1/R2: derivation

Let BB be the frozen teacher backbone, used as a feature extractor, and hϕh_\phi the trainable discriminator head, which pools the backbone’s intermediate features into one logit through small cross-attention branches with a single query token each. For fixed history, geometry, and timestep, the discriminator is Dϕ(x)=hϕ(B(x))D_\phi(x)=h_\phi(B(x)), where xx is the current video latent. R1 penalizes its input-gradient norm on real samples; R2 applies the same penalty on generated samples.

R1=Ex∼pdata∥∇xDϕ(x)∥22,R2=Ex∼pθ∥∇xDϕ(x)∥22.R_1=\mathbb{E}_{x\sim p_{\mathrm{data}}}\|\nabla_xD_\phi(x)\|_2^2, \quad R_2=\mathbb{E}_{x\sim p_\theta}\|\nabla_xD_\phi(x)\|_2^2.

Both penalties and their head-parameter gradients can be computed exactly without a double backward, because the backbone is frozen and the head computes its JVP with ordinary differentiable operations.

Gradient with a frozen backbone. Consider one sample and write z=B(x)z=B(x), A=JB(x)A=J_B(x), and u=∇zhϕ(z)u=\nabla_z h_\phi(z). A VJP gives the full input gradient g=A⊤ug=A^\top u, so the penalty value is simply its squared norm (step 1 in Figure 9). Next, a JVP carries that same gradient forward through the backbone, giving the feature direction v=Agv=Ag (step 2).

g=A⊤u,R=g⊤g,v=Ag.g=A^\top u, \qquad R=g^\top g, \qquad v=Ag.

Because the backbone is frozen, both its features zz and its input Jacobian AA are independent of the head parameters ϕ\phi. Differentiating the squared norm gives the following identity, with zz and vv held fixed during the head update and sg⁡\operatorname{sg} denoting stop-gradient.

∇ϕR=2(∂u∂ϕ)⊤v=∇ϕ ⁣[2u⊤sg⁡(v)]=∇ϕ ⁣[2Jzhϕ(z) sg⁡(v)].\nabla_\phi R=2\left(\frac{\partial u}{\partial\phi}\right)^\top v=\nabla_\phi\!\left[2u^\top\operatorname{sg}(v)\right]=\nabla_\phi\!\left[2J_z h_\phi(z)\,\operatorname{sg}(v)\right].

Here u⊤v=Jzhϕ(z) vu^\top v=J_z h_\phi(z)\,v is the directional derivative of the head along vv: a JVP of the head alone. The gradient of the penalty therefore needs only this one JVP, differentiated once with respect to ϕ\phi. The direction vv is recomputed at the current parameters on each update, then detached for this derivative. A nonlinear backbone is fully compatible with this identity: its Jacobian depends on xx, but not on ϕ\phi. Updating the backbone as well would introduce additional derivatives of its features and Jacobian, which this detached construction does not supply.

JVP in place of double backward. Backward-over-backward differentiates the computation that produced the input gradient. For a fused attention kernel, this requires either a differentiable backward or an additional implementation of its derivatives, with correct gradient paths through saved intermediates and incoming gradients. The regularizer therefore depends on the details of the backward kernel and its autograd bookkeeping.

A JVP follows the forward computation, carrying one tangent alongside each activation. Its per-operator rules compose in the same order as the network, so they can be implemented and checked one operator at a time.

For attention Y=PVY=PV, with P=softmax⁡(E)P=\operatorname{softmax}(E) and E=QK⊤/dE=QK^\top/\sqrt{d} (here VV is the value matrix, not the direction vv), the directional forward is analytic. A dot denotes a derivative along the supplied feature direction; the sum below runs over keys within each query row.

E˙=(Q˙K⊤+QK˙⊤)/d,P˙=P⊙(E˙−∑keysP⊙E˙),Y˙=P˙V+PV˙.\begin{aligned}\dot E&=(\dot QK^\top+Q\dot K^\top)/\sqrt d,\\\dot P&=P\odot\left(\dot E-\sum_{\mathrm{keys}}P\odot\dot E\right),\\\dot Y&=\dot P V+P\dot V.\end{aligned}

The frozen backbone needs only the resulting tangent values, which a fused kernel such as rCM’s FlashAttention JVP[21] can supply without a gradient graph. The head also needs gradients through its JVP, so it implements these expressions with matrix products, softmax and elementwise operations, and one ordinary backward supplies the mixed derivative with respect to features and head parameters. This is still a second-order derivative of the head, but no FlashAttention backward kernel is ever differentiated. Because each of the head’s cross-attention branches has a single query token, the explicit attention map grows only linearly with the number of feature tokens, which keeps the dense implementation practical.

The loss value and its gradient are assembled separately. Let s=2Jzhϕ(z) sg⁡(v)s=2J_z h_\phi(z)\,\operatorname{sg}(v). The expression below reports the full squared input-gradient norm as its value, while its gradient with respect to the head flows only through ss. It matches the penalty value and its first derivative with respect to ϕ\phi at the current update, not higher-order derivatives.

R~=sg⁡(∥g∥22)+s−sg⁡(s).\widetilde R=\operatorname{sg}(\|g\|_2^2)+s-\operatorname{sg}(s).

R1 and R2 apply this construction to real and generated inputs respectively, with generated inputs detached during the discriminator update; R~\widetilde R sums both, and the discriminator minimizes Lrel+γ2R~\mathcal{L}_{\mathrm{rel}}+\tfrac{\gamma}{2}\widetilde R (step 3).

Future Work

Depth and surface normals provide a practical conditioning interface, but dense geometry videos introduce redundancy. Once an object’s shape is known, its rigid motion can be described by a small set of pose parameters, while a condition video represents the resulting surfaces across many pixels and frames. Spatial downsampling reduces this cost, but can also remove thin structures, narrow gaps, and small contact changes needed for precise control.

Geometry alone also leaves some aspects of world state unspecified. A rotationally symmetric object can spin without changing its depth or normal maps, even as a painted marking rotates. Material, color, and object identity are likewise not uniquely determined by geometry. Visual history can help preserve these attributes, but may lose them after long occlusions or revisits beyond the memory window. These limitations call for conditions that communicate more of the state already maintained by the code world.

Future conditioning interfaces could use more abstract and compact representations of world state and dynamics, e.g., structured text describing object attributes and interaction states or high-dimensional latent features. The goal is to preserve physically relevant information with less redundancy than dense geometry videos. For example, structured text could specify the orientation and angular velocity of a bullet spinning around its long axis, even when its depth and normal maps remain unchanged.

Aligning these representations with the explicit state maintained by the code world and adapting pretrained video models to use them remain open challenges. A more expressive and efficient conditioning interface could support more precise control across a wider range of environments.

The interface also decides which worlds can be shown at all. A renderer that receives only geometry cannot tell a spinning wheel from a still one, so tasks that depend on such state are out of reach today. A richer interface widens the range of possible worlds, but someone still has to write them. We built the worlds in this post one at a time. Since a world is a program, the next step is to let coding agents write and revise worlds[1, 4], with the shared renderer supplying appearance. Generating a world is the easy part. The harder question is whether it deserves an agent’s time: whether the task can be solved from what the agent sees, whether a careless strategy fails, and whether there is something to learn that carries into the next round. We want these checks to run automatically, before any agent practices in a new world.

More worlds change how experience has to be kept. Here each playbook belongs to one world and is short enough to read in full. Across many worlds, agents must decide which lessons to keep, merge, or drop, and find the few that apply to the scene in front of them. Code worlds help: a world can be reset and replayed exactly, so a lesson can be tested before it is passed on. Lessons that hold in several worlds, such as how to confirm that a grasp held or which visual cues mislead, become general knowledge about acting in physical scenes, and unseen worlds can measure it.

The two grow together. Where agents fail tells us which worlds to write next, and each new world tests what they wrote down before.

References

  1. 01Wang et al. Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning. 2026. ↗
  2. 02Huang et al. SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models. 2026. ↗
  3. 03Zhou and Miao. Astronex-World 1.0: Real-Time Interactive World Model Foundation. 2026. ↗
  4. 04Chen et al. Code World Model: Coding Agent as World Brain. 2026. ↗
  5. 05Baker et al. Emergent Tool Use From Multi-Agent Autocurricula. 2019; ICLR 2020. ↗
  6. 06OpenAI. GPT-6 Astra. 2026. ↗
  7. 07NVIDIA. Cosmos 3: Omnimodal World Models for Physical AI. 2026. ↗
  8. 08Huang et al. ViPE: Video Pose Engine for 3D Geometric Perception. 2025. ↗
  9. 09Bin et al. NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors. ICCV 2025. ↗
  10. 10Huang et al. Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion. 2025. ↗
  11. 11Yin et al. One-step Diffusion with Distribution Matching Distillation. CVPR 2024. ↗
  12. 12Zhuang et al. Self Gradient Forcing: Native Long Video Extrapolation. 2026. ↗
  13. 13Guessous et al. FlexAttention: The Flexibility of PyTorch with the Performance of FlashAttention. PyTorch blog, 2024. ↗
  14. 14Zhu et al. Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation. 2026. ↗
  15. 15Zheng et al. Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models. 2026. ↗
  16. 16Zhao et al. Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout. 2026. ↗
  17. 17Yin et al. Improved Distribution Matching Distillation for Fast Image Synthesis. 2024. ↗
  18. 18Huang et al. The GAN is dead; long live the GAN! A Modern GAN Baseline. NeurIPS 2024. ↗
  19. 19Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. NeurIPS 2022. ↗
  20. 20Lin et al. Diffusion Adversarial Post-Training for One-Step Video Generation. 2025. ↗
  21. 21Zheng et al. Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency. rCM: FlashAttention JVP implementation. ↗

BibTeX

@misc{mirros2026evolvingagents,
    title  = {AgentGarten: Code Worlds for Evolving Agents},
    author = {{MirroS Team}},
    year   = {2026},
    month  = {Oct},
    url    = {https://mirros.ai/blog/worlds-for-evolving-agents},
    note   = {Blog post}
}