AI Summary

WorldCrafter nearly halves revisit error in video world models by borrowing a 3D model's brain for memory

WorldCrafter compresses a video world model's history with an encoder taken from a 3D novel-view synthesis model, then queries it with the upcoming camera path. When the camera returns to a scene it has already seen, LPIPS error drops from 0.487 (Lyra 2.0, the best of 8 baselines) to 0.255, and memory processing runs 21.7x faster than depth-warping.

Video world models forget where they’ve been. Walk the camera away from a room and back, and the couch has moved. WorldCrafter, from Peking University and Tencent’s ARC Lab, fixes a large chunk of that by compressing the model’s history with an encoder borrowed from a pretrained 3D novel-view synthesis model, then asking that memory “what should I see from here?” using the upcoming camera path. When the camera loops back to a scene it has already shown, WorldCrafter’s perceptual error (LPIPS) is 0.255, against 0.487 for the best of 8 recent world models.

  • Revisit consistency: LPIPS 0.487 to 0.255, PSNR 14.05 to 18.02 dB vs. Lyra 2.0, the strongest baseline, on 725 closed-loop camera trajectories of 528 to 1,648 frames
  • Memory processing: 0.062s per chunk vs. 1.346s for a depth-and-warp spatial memory, a 21.7x reduction, with no depth estimation at inference
  • Best camera control of the field: rotation error 13.5 vs. 16.1 (Lyra 2.0) and 23.5 (SANA-WM)
  • The distilled WorldCrafter-fast runs at 16 fps on 4 GPUs at 640x384, and scores even better on revisit consistency (LPIPS 0.186)
  • Consistency still breaks on long, complex trajectories, and history is re-encoded every chunk

Three ways to remember, and why each falls short

A video world model generates the next chunk of frames as the user moves the camera. To keep the world coherent, it needs memory beyond the last few seconds. The paper sorts prior approaches into three buckets:

  • Context memory: retrieve some old frames and let the model attend to them. Full-history attention is too expensive, so you pick a few frames, and whatever those frames didn’t see is lost.
  • Spatial memory: estimate depth, build geometry, warp old frames into the new viewpoint. This is only as good as the depth, and it struggles with anything that moves.
  • Implicit memory: learn a compressed representation of history. But representations trained for geometry tend to throw away the appearance detail you need to redraw the same couch with the same fabric.

WorldCrafter is an implicit memory, with a twist: the encoder comes from LagerNVS, a model already trained to render a scene from a new camera position given a few views. That task is exactly the question the world model needs answered on a revisit.

How it works

Per chunk of 9 latent frames at 640x384, the pipeline does four things:

  1. Retrieve history by coverage, not similarity. Pick the latest frame plus the old frames that together cover the most of the upcoming camera trajectory’s field of view. The point is complementary views, not the most similar ones.
  2. Encode. The LagerNVS-initialized memory encoder takes those VAE latents plus their camera poses (relative to the latest frame) and writes them into an implicit 3D-aware representation.
  3. Read out with the target poses. Instead of summarizing the whole representation into fixed tokens, the readout queries it with cameras sampled from where the user is about to go. The fixed memory-token budget (the size of 4 uncompressed frames) is spent only on what’s about to be on screen.
  4. Generate. The video DiT (initialized from Helios-base) denoises the new chunk conditioned on those memory tokens, a FramePack-style window of recent frames, the text prompt, and the target camera trajectory, which is injected as a relative-pose positional encoding (UCPE) in self-attention.

Training runs in four stages on 16 to 32 GPUs: fine-tune the DiT on 760K Open-Sora-Plan videos, add camera control, adapt the memory encoder to the video latent space, then train everything jointly. The fast variant is distilled with a pyramid scheme (3 resolutions, 2 denoising steps each).

The comparison against 8 recent world models on the authors’ benchmark (145 images, 5 metric camera trajectories each, all with closed-loop revisits):

Revisit error (LPIPS, lower is better)

Perceptual difference between what the model shows on a revisit and what it showed the first time. 725 videos.

LingBot-World 2
0.633
DreamX-World
0.627
Echo-WM
0.582
Alaya-EVOKE
0.565
SANA-WM
0.553
Matrix-Game 3.5
0.549
HY-WorldPlay
0.515
Lyra 2.0
0.487
WorldCrafter
0.255
WorldCrafter-fast
0.186

The 3D-consistency metric MEt3R tells the same story: 0.166 for WorldCrafter vs. 0.334 for Lyra 2.0 and 0.394 to 0.548 for the rest.

The eight baselines sit in a tight band; WorldCrafter isn’t close to them. On general visual quality (VBench) it doesn’t pay for that consistency either: its overall score of 81.91 is the highest, just ahead of Alaya-EVOKE (81.41) and SANA-WM (80.84), though SANA-WM wins on aesthetic quality.

It’s also much cheaper than warping

Spatial memory needs explicit geometry on every chunk. WorldCrafter reads the history latents directly. Per chunk, at 640x384:

Memory processing time per chunk

Spatial memory: Depth Anything 3 depth + alignment, then batched warping. WorldCrafter: memory encoding + readout.

Spatial memory
1.346s
WorldCrafter
0.062s

Spatial breakdown: 0.409s depth and alignment (dark), 0.937s warping (light). WorldCrafter: 0.049s encode, 0.013s readout. 21.7x less.

What each piece buys

The ablations (on the base model, before distillation) show every design choice carrying weight, and the biggest one is the memory itself. Swapping the 3D-aware encoder for plain context memory (just feeding 4 retrieved history frames in the same token budget) nearly doubles revisit error:

Ablations: revisit LPIPS (lower is better)

Each row changes one component of the full model.

Context memory instead
0.497
Pose-free readout
0.333
Frozen memory encoder
0.305
Similarity-based retrieval
0.296
Full WorldCrafter
0.255

Context memory also wrecks camera control (rotation error 26.5 vs. 13.5), which suggests the 3D-aware memory is doing double duty: it tells the model what’s there and where the camera is relative to it. Querying memory with the target poses beats a fixed summary, and fine-tuning the pretrained encoder jointly with the DiT beats freezing it. The generalizable lesson for anyone building memory into a generative system: retrieve for coverage, and read out with the question you’re about to ask.

Caveats

  • Self-built benchmark. All 725 test videos come from the authors’ own 145 images and trajectories. The baselines were run on it, but none of them were designed around it.
  • The distilled model beats the base model on memory, by a wide margin (LPIPS 0.186 vs. 0.255), and the paper doesn’t explain why. It does lose some camera accuracy and VBench quality in exchange.
  • Consistency still breaks on especially long or complex trajectories, per the authors.
  • History is re-encoded every chunk. Cheap relative to warping, but the same history gets processed again and again. A streaming encoder that folds in each new chunk is listed as future work.
  • 16 fps needs 4 GPUs, at 640x384. The base model’s speed isn’t reported.

Videos and supplementary material are on the project page. The takeaway is less about this particular model than the move it makes: a novel-view synthesis network already knows how to answer “what does this scene look like from over there,” and that is exactly the memory a world model needs.

#research #memory #world-models #video-generation

Liked this? Engineer's Codex sends one deep dive and a link roundup every week.

No spam. Unsubscribe anytime.