D-JEPA

Decision-aligned world modeling

A Decision-Aligned
Latent World Model

Predictive geometry describes futures. D-JEPA learns which relations among those futures produce better actions.

Built on pretrained predictive models, D-JEPA aligns candidate futures where actions are selected, integrates complementary predictive geometries through ordinal evidence, and writes the aligned structure back into future representations.

01 / THE QUESTION

A predicted future is useful
when it helps choose the action.

Predictive geometry and decision quality are not the same thing. The distance that makes a future look promising need not rank the candidate actions by their physical outcomes.

Our starting point

Learn the decision-relevant relations between candidate futures, then use those relations to align action selection.

The decision boundary is a learning target—not just a property inherited from pretraining.

LeWM and TD-JEPA rank correlation weakens as candidate shortlists narrow
Where prediction meets action selection. Ranking weakens within the small shortlists closest to execution. The diagnostic distinguishes average within-start correlations from pooled correlations over the same 96 starts.
02 / THE FRAMEWORK

From future prediction
to decision-aligned representation.

01

Decision alignment

Learn bounded, set-wise corrections from future–goal descriptors and ordinal structure. Restricted predictor adaptation offers a complementary source of decision evidence.

Explore the modules ↗
02

Multiple predictive geometries

Combine relational evidence from complementary predictors in one alignment framework, extending beyond a single predictive distance.

Read the protocols ↗
03

Representation realization

Encode a learned decision ordering in JEPA-compatible goal-distance geometry, and learn bounded, same-action updates across future time steps.

Explore checkpoints ↗
Two complementary representation mechanisms

PushT exact realization preserves earlier futures and encodes terminal ordering while retaining its embedded predictive and relational computation. Reacher learns same-action five-step future updates; its reported physical decisions use the relational selector. These experiments establish complementary mechanisms and do not imply a newly evaluated combined checkpoint or single-backbone distillation.

03 / RESULTS

Better future relations.
Better executed decisions.

Matched candidate sets isolate the quality of action selection across simulation, foundation models, autonomous driving and physical manipulation.

PUSHT · SUCCESS83.5987.89%

+4.30 pp over LeWM
Independent decision-boundary evaluation

ROBOTWIN · SUCCESS61.7276.76%

+15.04 pp over native VLA selection
Four language-conditioned tasks

REAL ROBOT · MACRO SUCCESS64.081.0%

+17.0 pp over V-JEPA 2-AC
PushT and two-cube stacking

DRIVING · PDMS57.3695.34

+37.98 over Drive-JEPA ranking
Seven difficult scenes

Core action-selection results
TaskEvaluationMatched referenceD-JEPAGain
PushTSuccess · 256 startsLeWM · 83.59%87.89%+4.30 pp
DMC-ReacherSuccess · 128 startsLeWM · 86.72%93.75%+7.03 pp
GranularStrict attainment · 64 startsDINO-WM · 7.81%18.75%+10.94 pp
Each task retains its own evaluation population and success criterion. Granular reports the strict Chamfer-distance threshold.
Cross-domain and generalization evidence
SettingMetric / scaleMatched referenceD-JEPAGain
Held-out PushObj shapesSuccess · 300 startsCalibrated fusion · 44.67%53.67%+9.00 pp
Appearance changesMean success · 7 conditionsCalibrated fusion · 62.86%73.14%+10.29 pp
RoboTwinStrict success · 4 tasksNative VLA · 61.72%76.76%+15.04 pp
Physical PiPERMacro success · 2 tasksV-JEPA 2-AC · 64.0%81.0%+17.0 pp
Autonomous drivingMean PDMS · 7 scenesDrive-JEPA · 57.3695.34+37.98
Appearance changes are evaluated on 50 new base starts per condition; RoboTwin uses 128 starts per task and PiPER 50 paired trials per task. Driving reports mean PDMS over seven difficult scenes.

Each row retains its own task metric and evaluation population. Full counts, paired outcomes and task-specific costs remain available in the paper and the decision-supervision dataset.

PushT independent 256-start success by method
PushT · independent 256
Reacher formal success by method
Reacher · independent 128
Granular success across distance thresholds
Granular · thresholds
Paired DINO-WM and D-JEPA Granular costs on the same starts
Granular · paired cost

TD denotes TD-JEPA; Ours denotes D-JEPA. Bars show success with 95% intervals, the curve varies the Granular success threshold, and each scatter point is one paired start. Click a panel to open the vector figure.

Transfer under geometric and visual change

Success on unseen I, small T and square shapes
Held-out shapes · success
Success across seven visual conditions
Appearance changes · success
Paired gains and losses on each held-out shape
Held-out shapes · paired outcomes
Net gains under each visual condition
Appearance changes · net gains

N: native selection; F: calibrated fusion; D: D-JEPA. Paired gains and losses compare D-JEPA with calibrated fusion on identical starts.

Contributions to decision alignment

Independent PushT success with plasticity, relational alignment and calibrated composition
Decision mechanisms · success
Matched single-source and dual-source predictive evidence
Predictive evidence · source ablation
Success rates and intervals for multiple predictive geometries
Multiple geometries · success
Median per-start inference latency of the three deployment configurations
Deployment · inference latency

The module and latency panels use the independent 256-start PushT population; Combined denotes calibrated composition. Source and multi-geometry ablations use the separate 128-start mechanism evaluation; Low and High denote the 95% interval bounds.

Decision mechanisms and native-distance deployment
ConfigurationDecision readoutSuccessMedian inference
Relational alignmentRelational score87.11%35.03 ms
Calibrated compositionGated composition87.89%52.16 ms
Ordinal realizationNative goal distance87.11%34.02 ms
Independent PushT evaluation, 256 starts. Ordinal realization recovers 100% of relational action choices and candidate ranks through native goal distance. Timing includes the predictive and relational computation, measured per start.

Decision structure in future representations

Success across shared candidate budgets
Candidate availability
Recovery of relational ranks through native distance
Exact ordinal recovery
Shared projection of future representations before and after realization
Future–goal geometry
Learned same-action residual coefficients across five future steps
Same-action temporal transport

The first three panels show PushT candidate-budget and ordinal-realization analyses; the final panel shows the separate Reacher five-step representation diagnostic. Budget bands span the minimum and maximum over 16 shared subset seeds.

Matched RoboTwin baseline failure and D-JEPA grasp success
RoboTwin. The same observation and candidate set produce different executed outcomes.
Matched driving candidates selected by Drive-JEPA and D-JEPA
Autonomous driving. Camera frames are shared logged observations; bird’s-eye views show the alternative simulated ego trajectories against logged traffic.
04 / IN MOTION

Same start. Same goal.
A different choice of action.

Curated baseline-failure / D-JEPA-success examples. These illustrate gains; aggregate results include gains, losses and ties.

PushT · start 0057

Temporal-Distance JEPA / D-JEPA

All 25 control actions of the fixed-horizon protocol, with initial context and terminal hold. This is not a receding-horizon full-task episode.

View the high-resolution timeline ↗

Reacher · start 0031

Temporal-Distance JEPA / D-JEPA

Full fixed 25-control-step sequence. Success is measured against the target joint configuration.

View the high-resolution timeline ↗

Granular · t00621-o10

DINO-WM / D-JEPA

All five actions and settling steps. High-resolution replay costs retain their own measured values; they do not replace formal aggregate measurements.

View the high-resolution timeline ↗
05 / BUILD ON D-JEPA

From a checkpoint
to a reproducible decision.

Task-specific checkpoints, decision supervision and a compact evaluation package.

Start with the independent PushT cached-decision replay, then reproduce the relational, composition and representation-realization paths from the same released protocol.

Full reproduction guide ↗
Run from the repository root
pip install -e '.[download]'
python scripts/download_artifacts.py --profile pusht-relational --dataset
python -m zipfile -e data/D-JEPA-supervision-v1.zip data
bash scripts/reproduce.sh

Task interfaces: robotic manipulation · autonomous driving · physical-robot workflows. Each guide describes inputs, training and evaluation commands, and environment dependencies.

The website reports completed evaluations only; release artifacts and protocol notes preserve the corresponding evidence identities.