Sparse Goal Steering
for Autoregressive Multi-Agent World Models
Replace a few motion tokens. Let the model continue.
Read the paperToken 101 → 192 changes the final half-second of the prefix. Control ends at 2 s; 31 of 32 saved continuations reach the goal by 5 s. Original figure paths, explanatory timing. No live inference.
Explore the rollout frame by frame
SGS
The first three motion tokens are identical. Each emits half a second of motion.
Compare the futures.
Figure 8 saved paths · SGS-repair · filtered reference. Selected earlier-study examples, including a failed intervention.
A few edits, a different route.
With at most two early token replacements, SGS-CEM improves guarded route success over native sampling.
Gain over native 10.79 pp, 95% paired interval [9.26, 12.36].
How SGS works
At selected early clocks, route geometry ranks eligible motion tokens. SGS retains a small support; the frozen model samples within it using its original logits and temperature.
A replacement is counted against the reference choice under the same history and shared noise. A requested mask can leave the token unchanged. Non-target agents use reference sampling throughout.
Search evaluates early joint prefixes through model continuations. The chosen prefix is fixed through 2 s, then tested with 32 fresh continuations excluded from search. Goal guidance ends at release.
Results and evaluation details
The primary confirmation uses 512 reserved source groups and 1,140 route requests. Controller selection uses a separate 64-group calibration set. Search is capped at 256 model predictions; groups receive equal weight.
Guarded success requires a forward goal-gate crossing by 5 s and target collision, offroad, validity, and dynamics checks through 8 s. These are target-level checks.
The separate development comparison uses 100 groups, 217 requests, and a filtered reference. Native achieves 28.85%, best-of-16 33.31%, and SGS-CEM 40.84%. These results use a different cohort and reference from the primary confirmation.
The experiments use SMART-family simulators. The paper discusses source independence, checkpoint exposure, and the difficulty of arbitrary fixed-time endpoints.
Original Figure 1 and animation details
The first three tokens are 455, 1360, and 1399 in both prefixes. The fourth changes from 101 to 192 during 1.5–2 s. The animation preserves the original unsmoothed vector paths and observed initial footprints. Its equal-scale frame is expanded slightly to show every saved 8 s path endpoint.
Token-selection graphics and playback timing are illustrative. The raw rollout files are absent from the local research snapshot, so this site renders paths already recorded in the paper. It adds no new samples or model predictions. The comparison slider uses three genuine native/SGS-repair pairs from the earlier eight-task cohort, separate from the primary confirmation.
Original vector figure ↗