Beckmann
World Models
Learn the destination.
Predict the future in one step.
Draft: September 25 · experiments below updated September 28
Robomimic Can
SPATIAL EPOCH 30per RGB chunk
spatial epoch 5 → 30
spatial epoch 5 → 30
From a long path
to a direct destination.
World models answer: what happens if I take these actions?
Repeated denoising makes each imagined future expensive.
Diffusion
Noise → repeated denoising → data.
DDPM ↗Few-step models
Learn larger jumps; trade steps for quality.
Consistency models ↗Drifting
Evolve generated samples during training. One call at test time.
Drifting ↗ DriftWorld ↗Beckmann maps
Learn the endpoint shared by an autonomous trajectory.
Beckmann Transport Models ↗Conceptual progression; distinct training objectives. Consistency models can use distillation or standalone training. BWM learns its generator without a pretrained generative teacher.
Many points.
One endpoint.
A terminal map stays constant along an autonomous transport trajectory.
Schematic geometry · transport time is separate from video time.
Population identities under BTM’s assumptions; they motivate learning, not a guarantee for the trained video model.
Mix noise with a recorded future.
Differentiate only the future; hold context fixed.
Regress to the stopped target; anchor clean and near-data inputs.
sg = stop-gradient · η = 0.25 · transport uses detached error weighting and output-layer gradient balancing. Sampled chords are not flow trajectories.
Supervise the future
you actually generate.
The deployment query is pure noise + history + actions.
Put the learning signal directly on that endpoint.








Train where you sample.
Coordinate-mean squared error on the pure-noise prediction; a clean boundary alone is insufficient.
Focus on what changes.
Target motion masks, within-chunk differences, and short detached-history recovery.
Preserve spatial detail.
VGG16 at three scales: normalized channels, averaged locations. Gradients reach the generated image; training only.
Staged training; component effects are not isolated. Spatial weight ramps to 0.05; RGB feedback remains unclipped.
Source MSE can suppress diversity: prediction accuracy does not establish calibrated sampling. Pipeline arrows illustrate controls; frames are saved predictions.
Endpoint supervision improves the complete recipe. Added complexity alone does not: Can’s learned-direction branch raised LPIPS from 0.0225 → 0.0854.
More training.
Still one call.
Completed 30-epoch spatial continuations on PushT and Can.
“30 epochs” counts the spatial stage. Including inherited training: 301,028 updates on PushT; 45,000 on Can. Curves show recorded precision; tables use final server evaluations.
PushT
Held-out predictions · same task split| Model / checkpoint | Calls / chunk | MSE ↓ | Moving MSE ↓ | LPIPS ↓ | SSIM ↑ | Full-episode MSE ↓ |
|---|---|---|---|---|---|---|
| BWM + spatialspatial epoch 30 · final | 1 | 0.0098 | 0.1247 | 0.0280 | 0.975 | 0.0154 |
| BWM + spatialspatial epoch 5 | 1 | 0.0155 | 0.1933 | 0.0407 | 0.964 | 0.0231 |
| BWMrobust recipe | 1 | 0.0199 | 0.2419 | 0.0564 | 0.953 | 0.0299 |
| DriftWorld5-epoch baseline | 1 | 0.0314 | 0.4093 | 0.0790 | 0.927 | 0.0451 |
| MSE U-Net5-epoch baseline | 1 | 0.0303 | 0.4008 | 0.0804 | 0.943 | 0.0427 |
| GPC diffusion5-epoch baseline | 12 | 0.0630 | 1.4241 | 0.1609 | 0.904 | 0.0753 |
| AVDC diffusion5-epoch baseline | 100 | 0.0684 | 1.1840 | 0.1860 | 0.895 | 0.0795 |
Robomimic Can
Held-out predictions · same task split| Model / checkpoint | Calls / chunk | MSE ↓ | Moving MSE ↓ | LPIPS ↓ | SSIM ↑ | Full-episode MSE ↓ |
|---|---|---|---|---|---|---|
| BWM + spatialspatial epoch 30 · final | 1 | 0.0022 | 0.0242 | 0.0056 | 0.984 | 0.0082 |
| BWM + spatialspatial epoch 5 | 1 | 0.0028 | 0.0298 | 0.0073 | 0.979 | 0.0114 |
| BWMrobust recipe | 1 | 0.0031 | 0.0316 | 0.0090 | 0.974 | 0.0112 |
| DriftWorld5-epoch baseline | 1 | 0.0104 | 0.0769 | 0.0682 | 0.856 | 0.0256 |
| MSE U-Net5-epoch baseline | 1 | 0.0051 | 0.0544 | 0.0104 | 0.969 | 0.0138 |
| GPC diffusion5-epoch baseline | 6 | 0.0276 | 0.3103 | 0.0589 | 0.910 | 0.0380 |
| AVDC diffusion5-epoch baseline | 100 | 0.0400 | 0.2644 | 0.1194 | 0.801 | 0.1768 |
64 frames = 4 real context + 60 generated. PushT: 388 eligible videos / 400 full episodes; Can: 30 / 30. MSE on RGB [−1, 1].
BWM rows have cumulative training; local baselines use 5 epochs. These are not equal-total-compute comparisons.
One call buys real throughput.
A800 · batch 1 · model onlyDedicated September 26 timing at spatial epoch 5 · FP32, TF32 off · 4 frames/chunk PushT, 2 Can. Later checkpoints retain the same one-call sampler; these speed measurements were not rerun at epoch 30.
One scene.
Every method.
Synchronized comparisons from the earlier spatial-epoch-5 evaluation.
The epoch-30 examples are at the top of the page.
Video could not load. Please select another example.
All eight saved examples available. Can AVDC has no final-checkpoint clip. Video playback is 10 fps, not inference speed. Selected examples do not replace the complete-cohort metrics above.
The same idea,
in latent space.
Bridge-V2 and RT-1 · native 256px video · SD3 VAE latents.
Native DriftWorld U-Net: 175.4M / 160.3M parameters · one generator call per frame · temporal loss is zero for a one-frame endpoint.
Completed native evaluations
Ours + spatial only · subset cohort| Dataset | SSIM ↑ | PSNR ↑ | LPIPS ↓ | FID ↓ | FVD ↓ | Calls / frame |
|---|---|---|---|---|---|---|
| Bridge-V2subset · epoch 30 | 0.796 | 22.13 | 0.155 | 68.86 | 592.24 | 1 |
| RT-1subset · epoch 30 | 0.802 | 23.15 | 0.166 | 75.53 | 535.59 | 1 |
Completed subset runs: 2,078 / 2,041 eligible training episodes; 215 / 198 validation videos. Eight-frame chunks re-anchor on ground truth; metrics compare against VAE-reconstructed ground truth. FVD uses 201 / 193 videos. Different cohort and training data from published full-data baselines.
Bridge-V2: 52,969 training episodes · RT-1: 86,634. Full-data runs active; matched-subset baselines queued at the September 29, 00:53 UTC check.
Subset generator latency: 16.3 ms/frame Bridge; 13.1 ms/frame RT-1. VAE decode adds about 26 ms/frame. Published full-data scores are not a controlled comparison with this cohort.
Prediction improves.
Planning stays close.
PushT’s final GPC-RANK difference includes zero.
Absolute policy-outcome prediction remains accurate.
| Model | Planning ↑ | Policy r ↑ | IoU MAE ↓ |
|---|---|---|---|
| BWM + spatialspatial epoch 30 | 0.721 | 0.928 | 0.010 |
| BWM + spatialspatial epoch 5 | 0.708 | 0.946 | 0.010 |
| DriftWorld5-epoch baseline | 0.709 | 0.857 | 0.031 |
| MSE U-Net5-epoch baseline | 0.695 | 0.915 | 0.012 |
| GPC diffusion5-epoch baseline | 0.615 | 0.366 | 0.513 |
Planning: 50 fixed evaluation seeds. Policy: 7 policies × 300 paired trials. Stage-30 values rounded from the committed evaluation report.
Actions matter to the prediction.
Spatial epoch 5 · one-chunk probe64 held-out windows with real history. Swapped-action error measures sensitivity, not accuracy against real counterfactual outcomes. Better video and action sensitivity do not establish better planning.
Conditional diversity
Does accuracy preserve plausible alternatives?
Action consequences
Can better endpoints reduce real selection regret?
Full-data scaling
Do gains persist on matched native cohorts?