Fast LeWorldModel

Yuntian Gao, Xiangyu Xu†

Xi'an Jiaotong University  |  †Corresponding author

Open-Loop Rollout Visualization

Two-Room

start

Two-Room open-loop rollout start frame

goal

Two-Room open-loop rollout goal frame

LeWM

Ready

Fast-LeWM

Ready

Reacher

start

Reacher open-loop rollout start frame

goal

Reacher open-loop rollout goal frame

LeWM

Ready

Fast-LeWM

Ready

Planning

Cube

LeWM

Sequential rollout

CEM ready
Rollout in Planning...
GT
LeWM
Fast-LeWM

Parallel prefix

CEM ready
prefix pred Rollout in Planning...
GT
Fast-LeWM

PushT

LeWM

Sequential rollout

CEM ready
Rollout in Planning...
GT
LeWM
Fast-LeWM

Parallel prefix

CEM ready
prefix pred Rollout in Planning...
GT
Fast-LeWM

Abstract

Joint-Embedding Predictive Architectures (JEPAs), including LeWorldModel (LeWM), are promising reconstruction-free visual world models. However, LeWM evaluates candidate action sequences through repeated one-step latent transitions, which makes planning slow and allows latent prediction errors to accumulate over long horizons.

Fast-LeWM replaces repeated local rollout with action-prefix prediction. Given the current latent and a candidate action sequence, it encodes prefixes of that sequence and predicts the future latents reached after executing those prefixes in parallel. Joint multi-horizon training teaches the model how states evolve under different action prefixes. During planning, each future latent can be evaluated directly from its prefix without rolling through intermediate imagined states. Across four tasks, Fast-LeWM improves average planning success from 85.8% to 90.5% while reducing CEM solve time from 33.7s to 16.1s. It also lowers open-loop latent prediction loss and slows its growth over longer horizons.

Method: action-prefix prediction

Using prefixes of the candidate action sequence as multi-horizon queries, Fast-LeWM predicts future latents in parallel from the observed anchor latent.

Fast-LeWM training pipeline with visual encoder, causal action-prefix encoder, and parallel latent predictor
Training pipeline. A state token and causally masked action tokens produce prefix tokens; the parallel predictor learns every future latent horizon through dense supervision.

Results

Fast-LeWM is evaluated on the same goal-conditioned planning tasks and protocol as LeWM: Two-Room, Reacher, PushT, and OGBench-Cube.

4.8x

faster dynamics evaluation: 19.7s to 4.1s.

52.2%

lower full CEM solve time: 33.7s to 16.1s.

90.5%

average success rate, improved from LeWM's 85.8%.

Dynamics evaluation time falls from 19.7 to 4.1 seconds; full CEM planning is faster with higher success across four tasks
Planning efficiency and success. Parallel prefix prediction reduces dynamics evaluation and full CEM solve time while improving average success across the four tasks.

Planning success (%)

MethodTwo-RoomReacherPushTCubeAvg.
PLDM9778786579.5
DINO-WM10079748684.8
LeWM8786967485.8
Fast-LeWM9888968090.5
Fast-LeWM + Self-Consistency9890988292.0
Open-loop latent prediction loss over 50 environment steps across Two-Room, Reacher, PushT, and Cube
Open-loop latent prediction over 50 environment steps. Fast-LeWM has lower loss at every evaluated step and a smaller fitted error-growth slope on all four tasks.

BibTeX

@misc{gao2026fastleworldmodel,
      title={Fast LeWorldModel}, 
      author={Yuntian Gao and Xiangyu Xu},
      year={2026},
      eprint={2606.26217},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2606.26217v2},
}