Long spans for supervision, not for attention. Memorizon trains on minutes-long spans that contain both visits to a place, while each chunk attends only to a small bank of retrieved frames, so the cost stays bounded.
The bank is bounded by kK: it levels off at every K while the candidate pool keeps growing, from 36 to 1556 latents.
So the sequence the model reads stays short however long the span: going from 100 s to 400 s adds only 12% to the step time.
The 4-step model on more web photographs, 60 s each. Each camera path is planned from the photograph's depth so it stays in the free space it shows, and comes back to the first view within 30 s. Hover to play; click to open full size.