Crowd simulation · world models · 2026

Controllable Crowd Generation through World-Model Planning

Ctrl-CWM learns how people move from real pedestrian video, then plans every step in imagination, so a new objective can steer a running crowd without retraining.

1Yonsei University2GIST*Corresponding author
Without controlWith control · Ctrl-CWM

A collision mid-scene. Both sides are the same Ctrl-CWM rollout. At the moment of impact, only the right side is given an avoidance objective, with no retraining.

Abstract

Crowds that adapt when the world changes

Crowd simulation plays a central role in robot navigation, autonomous driving, and urban planning. For these applications, realistic simulation requires crowds to adapt their behavior to environmental changes and user objectives. However, existing methods that rely on predefined control settings have limited flexibility in accommodating new user-specified objectives. To address this limitation, we propose Ctrl-CWM, a multi-agent Controllable Crowd World Model that integrates crowd generation and run-time control. Our key idea is to adapt the world-model principle of planning using imagined futures to crowd simulation. To this end, Ctrl-CWM consists of an encoder that learns a representation of human motion dynamics, an actor that proposes pedestrian displacements, a critic that evaluates imagined crowd trajectories, and a planner that selects actions. We first learn human motion dynamics through trajectory prediction on real-world pedestrian videos and then freeze the encoder to preserve them. Using this representation, the actor generates imagined crowd trajectories through repeated state updates, and the planner combines the critic's scores with user costs to select actions. Repeated planning advances the simulated crowd, while additional user costs introduce new control objectives without retraining. We extensively evaluate crowd generation under varied agent arrival conditions and run-time control across avoidance and attraction scenarios. Ctrl-CWM outperforms the state-of-the-art method on most crowd realism and collision metrics, and adapts crowd behaviors to user-specified objectives introduced during simulation.

−22% DTWvs. CrowdES, average of 7 benchmarks (random-surface arrivals)
9 / 16average scores ranked first, second on the rest
0.950avoidance compliance, multiple zones (CrowdES 0.579)
~10 msper planning call (K = 16, I = 4, H = 4)
Method

Imagine, rank, then act

Pedestrian videos record how people moved, but they define no actions and no rewards. Ctrl-CWM builds both from trajectory prediction and then plans over imagined crowd futures.

Overview: learning the world model from prediction, learning behavior in imagination, and planning and control with a user command
Trajectory prediction trains the state encoder hθ on real pedestrian data, which is then frozen. The actor πψ and critic Vφ learn over imagined rollouts on this representation, and CEM adds a user cost to the critic score for run-time control.
  1. World model from prediction

    A scene U-Net and an interaction encoder (attention and a GRU, fused by FiLM) learn an agent-centric state through goal, waypoint, and one-step heatmap prediction on real trajectories. The encoder is then frozen to keep the learned human dynamics.

  2. Behavior in imagination

    The actor learns one-step displacements through 12-step imagined rollouts with soft-DTW, kinematic, and adversarial losses. The critic learns to rank imagined futures by realism and collisions with a listwise objective.

  3. Planning and control

    At every step, CEM samples displacements around the actor's proposal, rolls each forward in imagination, and scores it. A user cost steers the crowd; with the cost at zero the same planner generates crowds.

J = Vφ(ŝt+1:t+H) − λuser · cuser(ŝt+1:t+H) The critic keeps motion human-like and the user cost encodes the objective, such as avoiding a zone or approaching an attraction.
Results

Change the objective while the crowd is moving

Without control · CrowdESWith control · Ctrl-CWM
Objective
Scene

Results

When something happens mid-scene

Without controlWith control · Ctrl-CWM
Scene

Results

Through a pedestrian's eyes

Third-person
WithoutWith control
Ego-centric
WithoutWith control
Results

Quantitative results

Crowd generation · averageDens.Freq.Cov.Pop.Kinem.DTWDiv.Col.(%)
Random-surface arrivals
ORCA1.0440.0990.0992.6591.0152.8570.1960.009
CrowdES0.7460.1280.1281.8710.6052.2040.2943.448
Ctrl-CWM (ours)0.4770.1110.1111.2440.5451.7130.2851.739
Diffusion arrivals
ORCA1.0290.1060.1062.5951.1983.2280.1800.065
CrowdES0.1470.0520.0510.3150.5252.1260.2971.930
Ctrl-CWM (ours)0.1100.0460.0460.3720.6601.7380.3101.572

Averages over five ETH–UCY folds, SDD, and GCS (Table 1 of the paper). ORCA optimizes collision avoidance directly and has the lowest collision rate, but the highest density error and DTW.

Avoidance compliance ↑ETHHOTELUNIVZARA1ZARA2SDDGCSAVG
Disc zone
ORCA + obstacle1.000−1.1540.5300.4281.0000.3440.9990.450
CrowdES + map0.5990.5460.7270.4690.5980.3170.4360.527
Ctrl-CWM (ours)0.8370.9880.8310.7520.7140.8770.8270.832
Rectangle zone
ORCA + obstacle1.000−1.4840.526−0.0491.000−1.8270.595−0.034
CrowdES + map0.8830.8080.7790.7900.8240.1140.3010.643
Ctrl-CWM (ours)0.8960.8810.6830.9880.9120.9340.9320.889
Multiple zones
ORCA + obstacle0.9730.9990.601−1.952−0.304−0.4730.7680.087
CrowdES + map0.4790.5420.5490.7810.7510.5250.4240.579
Ctrl-CWM (ours)0.9920.9970.9930.9830.9710.9080.8080.950

Compliance Cavoid = 1 − occcmd/occfree: one means the controlled crowd never enters the zone, zero means no reduction. Ctrl-CWM receives the objective online; the baselines get the obstacle from initialization.

Cite

BibTeX

@article{lee2026ctrlcwm,
  title   = {Controllable Crowd Generation through World-Model Planning},
  author  = {Lee, JunGyu and Shin, Jisu and Shin, Seunghyun and Jeon, Hae-Gon},
  journal = {arXiv preprint arXiv:2610.09438},
  year    = {2026}
}