AR
arXiv CS.AI
7/9/2026

Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix
Short summary
Researchers identify an instruction leakage trap in goal-conditioned compact world models: a predictor achieves 0.90 relation-readout accuracy by transcribing the instruction rather than perceiving the scene. Withholding the goal collapses accuracy to chance (0.27), and counterfactual instructions make the model follow false instructions 94.5% of the time. The fix is to keep the goal out of the dynamics model and supervise the read path, recovering instruction-independent grounding at 0.88.
- •Goal-conditioned world models can fake spatial grounding by transcribing instructions
- •Withholding the goal collapses accuracy from 0.90 to 0.27, proving instruction leakage
- •Fix: keep the goal in the planner's cost, not the dynamics model, to recover genuine grounding
Generated with AI, which can make mistakes.
Is this a good recommendation for you?