WORLD MODELING · SEPTEMBER 2026

Beckmann
World Models

Learn the destination.
Predict the future in one step.

Adam LeeBerkeley
Shobhit AgarwalStanford
JP SuhGoogle DeepMind

Draft: September 25 · experiments below updated September 28

PushT

SPATIAL EPOCH 30
Ground truthBWM + spatial
EXAMPLE

Robomimic Can

SPATIAL EPOCH 30
Ground truthBWM + spatial
EXAMPLE
1generator call
per RGB chunk
37%lower PushT MSE
spatial epoch 5 → 30
22%lower Can MSE
spatial epoch 5 → 30
01 / THE IDEA

From a long path
to a direct destination.

World models answer: what happens if I take these actions?
Repeated denoising makes each imagined future expensive.

ITERATIVE GENERATION

Diffusion

Noise → repeated denoising → data.

DDPM ↗

Conceptual progression; distinct training objectives. Consistency models can use distillation or standalone training. BWM learns its generator without a pretrained generative teacher.

02 / TRANSPORT GEOMETRY

Many points.
One endpoint.

A terminal map stays constant along an autonomous transport trajectory.

FIX THE HISTORY AND ACTIONS Noise / sourceData manifoldyxT(x) = y

Schematic geometry · transport time is separate from video time.

AUTONOMOUS FLOW
X˙τ=bc(Xτ)\dot X_\tau=b_c(X_\tau)
SAME DESTINATION
JxTc(x) bc(x)=0J_xT_c(x)\,b_c(x)=0
FIX THE DATA
Tc(y)=yT_c(y)=y

Population identities under BTM’s assumptions; they motivate learning, not a guarantee for the trained video model.

01 · SAMPLE A TRAINING CHORD
xs=(1−s)z+sy,s∼U[0,1]x_s=(1-s)z+sy,\quad s\sim\mathcal U[0,1]

Mix noise with a recorded future.

02 · FORM A STOPPED TARGET
u=sg⁡ ⁣[Tθ(xs,c)+ηJxTθ(xs,c)(y−z)]u=\operatorname{sg}\!\left[T_\theta(x_s,c)+\eta J_xT_\theta(x_s,c)(y-z)\right]

Differentiate only the future; hold context fixed.

03 · LEARN THE MAP
LBTM=λLtr+Lboundary+Lnear\mathcal L_{\mathrm{BTM}}=\lambda\mathcal L_{\mathrm{tr}}+\mathcal L_{\mathrm{boundary}}+\mathcal L_{\mathrm{near}}

Regress to the stopped target; anchor clean and near-data inputs.

sg = stop-gradient · η = 0.25 · transport uses detached error weighting and output-layer gradient balancing. Sampled chords are not flow trajectories.

03 / THE WORLD MODEL

Supervise the future
you actually generate.

The deployment query is pure noise + history + actions.
Put the learning signal directly on that endpoint.

CONDITION c
Observed PushT frame oneObserved PushT frame twoObserved PushT frame threeObserved PushT frame four
4 observed frames
↘↓↙←
proposed actions
NOISE z
Tθ(z,h,a)T_\theta(z,h,a)
Residual U-Net · one call
GENERATED ENDPOINT
Generated future frame oneGenerated future frame twoGenerated future frame threeGenerated future frame four
Future chunk · feed back to predict again
THE ENTIRE RGB SAMPLER
y^=Tθ(z,h,a)=z+Uθ(z,h,a),z∼N(0,1.352I)\widehat y=T_\theta(z,h,a)=z+U_\theta(z,h,a),\qquad z\sim\mathcal N(0,1.35^2I)
PushT: 4 frames/call · Can: 2 frames/call · no interpolation-time input
A / SOURCE ANCHOR

Train where you sample.

Lsrc=E∥Tθ(z,c)−y∥D2\mathcal L_{\mathrm{src}}=\mathbb E\|T_\theta(z,c)-y\|_D^2

Coordinate-mean squared error on the pure-noise prediction; a clean boundary alone is insufficient.

B / MOTION + CONTEXT

Focus on what changes.

0.1 motion0.05 temporal0.01 range

Target motion masks, within-chunk differences, and short detached-history recovery.

C / SPATIAL FEATURES

Preserve spatial detail.

Lsp=13∑ℓ∥ϕˉℓ(y^)−ϕˉℓ(y)∥2\mathcal L_{\mathrm{sp}}=\frac{1}{3}\sum_\ell\|\bar\phi_\ell(\widehat y)-\bar\phi_\ell(y)\|^2

VGG16 at three scales: normalized channels, averaged locations. Gradients reach the generated image; training only.

Transport + boundary→+ source→+ motion / recovery→+ spatial

Staged training; component effects are not isolated. Spatial weight ramps to 0.05; RGB feedback remains unclipped.

Source MSE can suppress diversity: prediction accuracy does not establish calibrated sampling. Pipeline arrows illustrate controls; frames are saved predictions.

PushTRecipe endpoints · LPIPS ↓
PushT recipe endpoint LPIPS decreases from 0.0746 with source supervision to 0.0564 with recovery, 0.0407 with spatial epoch 5 and 0.0280 at spatial epoch 30. Training is cumulative.
Robomimic CanRecipe endpoints · LPIPS ↓
Can recipe endpoint LPIPS decreases from 0.0225 with source supervision to 0.0090 with recovery, 0.0073 with spatial epoch 5 and 0.0056 at spatial epoch 30. Training is cumulative.
WHAT WE LEARNED

Endpoint supervision improves the complete recipe. Added complexity alone does not: Can’s learned-direction branch raised LPIPS from 0.0225 → 0.0854.

04 / NEW RESULTS · SEPTEMBER 28

More training.
Still one call.

Completed 30-epoch spatial continuations on PushT and Can.

PushTMSE 0.0155 → 0.0098
PushT MSE across spatial-stage epochs 1 through 30; epoch 5 0.0155, final 0.0098, best at epoch 24 0.0095.
Robomimic CanMSE 0.0028 → 0.0022
Can MSE across spatial-stage epochs 1 through 30; epoch 5 0.0028 and final 0.0022.

“30 epochs” counts the spatial stage. Including inherited training: 301,028 updates on PushT; 45,000 on Can. Curves show recorded precision; tables use final server evaluations.

PushT

Held-out predictions · same task split
Model / checkpointCalls / chunkMSE ↓Moving MSE ↓LPIPS ↓SSIM ↑Full-episode MSE ↓
BWM + spatialspatial epoch 30 · final10.00980.12470.02800.9750.0154
BWM + spatialspatial epoch 510.01550.19330.04070.9640.0231
BWMrobust recipe10.01990.24190.05640.9530.0299
DriftWorld5-epoch baseline10.03140.40930.07900.9270.0451
MSE U-Net5-epoch baseline10.03030.40080.08040.9430.0427
GPC diffusion5-epoch baseline120.06301.42410.16090.9040.0753
AVDC diffusion5-epoch baseline1000.06841.18400.18600.8950.0795

Robomimic Can

Held-out predictions · same task split
Model / checkpointCalls / chunkMSE ↓Moving MSE ↓LPIPS ↓SSIM ↑Full-episode MSE ↓
BWM + spatialspatial epoch 30 · final10.00220.02420.00560.9840.0082
BWM + spatialspatial epoch 510.00280.02980.00730.9790.0114
BWMrobust recipe10.00310.03160.00900.9740.0112
DriftWorld5-epoch baseline10.01040.07690.06820.8560.0256
MSE U-Net5-epoch baseline10.00510.05440.01040.9690.0138
GPC diffusion5-epoch baseline60.02760.31030.05890.9100.0380
AVDC diffusion5-epoch baseline1000.04000.26440.11940.8010.1768

64 frames = 4 real context + 60 generated. PushT: 388 eligible videos / 400 full episodes; Can: 30 / 30. MSE on RGB [−1, 1].

BWM rows have cumulative training; local baselines use 5 epochs. These are not equal-total-compute comparisons.

Long-trained referenceReleased DriftWorld, PushT: MSE 0.0085 · LPIPS 0.0163 · 1.18M steps.Still better perceptual quality; evaluation-episode exclusion is unverified.

One call buys real throughput.

A800 · batch 1 · model only
PushT6.8× GPC · 101× AVDC
Stage-five PushT measured throughput comparison: BWM plus spatial 193.9 frames per second, GPC 28.4, AVDC 1.9.
Robomimic Can2.9× GPC · 99× AVDC
Stage-five Can measured throughput comparison: BWM plus spatial 83.8 frames per second, GPC 28.4, AVDC 0.8.

Dedicated September 26 timing at spatial epoch 5 · FP32, TF32 off · 4 frames/chunk PushT, 2 Can. Later checkpoints retain the same one-call sampler; these speed measurements were not rerun at epoch 30.

05 / LOOK AT THE PREDICTIONS

One scene.
Every method.

Synchronized comparisons from the earlier spatial-epoch-5 evaluation.
The epoch-30 examples are at the top of the page.

01 / 64

All eight saved examples available. Can AVDC has no final-checkpoint clip. Video playback is 10 fps, not inference speed. Selected examples do not replace the complete-cohort metrics above.

06 / BEYOND SIMULATION · NEW

The same idea,
in latent space.

Bridge-V2 and RT-1 · native 256px video · SD3 VAE latents.

4 context frames→Frozen VAE→Conditional endpoint map→Decode future frame

Native DriftWorld U-Net: 175.4M / 160.3M parameters · one generator call per frame · temporal loss is zero for a one-frame endpoint.

Perceptual error fallsLPIPS ↓
Native subset LPIPS at epochs 10, 20, 30: Bridge 0.493, 0.183, 0.155; RT-1 0.529, 0.196, 0.166.
Structure improvesSSIM ↑
Native subset SSIM at epochs 10, 20, 30: Bridge 0.417, 0.750, 0.796; RT-1 0.387, 0.762, 0.802.

Completed native evaluations

Ours + spatial only · subset cohort
DatasetSSIM ↑PSNR ↑LPIPS ↓FID ↓FVD ↓Calls / frame
Bridge-V2subset · epoch 300.79622.130.15568.86592.241
RT-1subset · epoch 300.80223.150.16675.53535.591

Completed subset runs: 2,078 / 2,041 eligible training episodes; 215 / 198 validation videos. Eight-frame chunks re-anchor on ground truth; metrics compare against VAE-reconstructed ground truth. FVD uses 201 / 193 videos. Different cohort and training data from published full-data baselines.

Full-data extension: results pending

Bridge-V2: 52,969 training episodes · RT-1: 86,634. Full-data runs active; matched-subset baselines queued at the September 29, 00:53 UTC check.

Latest run table ↗

Subset generator latency: 16.3 ms/frame Bridge; 13.1 ms/frame RT-1. VAE decode adds about 26 ms/frame. Published full-data scores are not a controlled comparison with this cohort.

07 / WHAT BETTER VIDEO DOES—AND DOESN’T—BUY

Prediction improves.
Planning stays close.

PushT’s final GPC-RANK difference includes zero.
Absolute policy-outcome prediction remains accurate.

Planning vs. DriftWorldPaired 95% intervals
Final epoch-30 BWM planning difference is plus 0.012 with 95 percent interval minus 0.047 to plus 0.071; the interval includes zero.
ModelPlanning ↑Policy r ↑IoU MAE ↓
BWM + spatialspatial epoch 300.7210.9280.010
BWM + spatialspatial epoch 50.7080.9460.010
DriftWorld5-epoch baseline0.7090.8570.031
MSE U-Net5-epoch baseline0.6950.9150.012
GPC diffusion5-epoch baseline0.6150.3660.513

Planning: 50 fixed evaluation seeds. Policy: 7 policies × 300 paired trials. Stage-30 values rounded from the committed evaluation report.

Actions matter to the prediction.

Spatial epoch 5 · one-chunk probe
PushT11.0× error with swapped actions
PushT source endpoint moving-region error increases from 0.0642 to 0.7089 when actions are swapped for BWM plus spatial.
Robomimic Can6.2× error with swapped actions
Can source endpoint moving-region error increases from 0.0073 to 0.0451 when actions are swapped for BWM plus spatial.

64 held-out windows with real history. Swapped-action error measures sensitivity, not accuracy against real counterfactual outcomes. Better video and action sensitivity do not establish better planning.

STILL OPEN

Conditional diversity
Does accuracy preserve plausible alternatives?

Action consequences
Can better endpoints reduce real selection regret?

Full-data scaling
Do gains persist on matched native cohorts?