Memorizon: Training World Models Beyond Their Context Window

Institute of Foundation Models (IFM), MBZUAI
arXivsoon Code Weights
TL;DR

Long spans for supervision, not for attention. Memorizon trains on minutes-long spans that contain both visits to a place, while each chunk attends only to a small bank of retrieved frames, so the cost stays bounded.

Training Long Span at Bounded Cost

first visit the camera returns one episode ordinary training A T1⋯Tk ✗ first visit is outside the window sampled training span A history (not tokenised) R T1⋯Tk union of per-chunk top-K → shared bank model input A B R T1⋯Tk loss on T only
A first frame B bank R recent T target chunks
  1. A place is often revisited minutes later, but ordinary clip training only sees a short tail of the episode, so no sample ever contains both visits.
  2. Memorizon samples a long span that holds both visits, yet computes the loss only on its last k chunks.
  3. Each scored chunk retrieves its own top-K latents by camera co-visibility; their union forms one shared bank, so the model reads A | B | R | T  and the sequence stays bounded however long the span is.
beyond 100 s 10 100 1000 4110 s 10125 s 20150 s 401100 s 801200 s 1601400 s training span (latent frames) bank size |B| bank empty candidate pool K = 2 K = 4 K = 6 K = 8 K = 10

The bank is bounded by kK: it levels off at every K while the candidate pool keeps growing, from 36 to 1556 latents.

So the sequence the model reads stays short however long the span: going from 100 s to 400 s adds only 12% to the step time.

Ablation Study