Skip to content

Benchmark Overview

Grounded in our operational definition of world models and the formal criterion introduced in Sec. 2, WorldAtlas formulates world generation as a set of conditional next-world generation tasks across video, 3D, and 4D modalities.

Following a next-scene evaluation perspective, each task instance is specified by a tuple (C,N,L)(C, N, L):

  • C={X,P}C = \{X, P\} — current evidence: an image or video observation XX and an optional textual description PP
  • NN — target world specification: the next scene, dynamic change, object transformation, or interaction outcome to be generated
  • LL — layout and control specification: camera trajectory, viewpoint change, navigation path, object-level control, action instruction, or fixed-camera constraint

Given this specification, a model is required to generate an output artifact

O=gworld(wproc(C,N,L)),O = g_{\mathrm{world}}(w_{\mathrm{proc}}(C, N, L)),

where:

  • gworldg_{\mathrm{world}} denotes the evaluated generative system
  • wprocw_{\mathrm{proc}} denotes model-specific preprocessing
  • OO may be a video sequence, a 3D scene, a 4D dynamic asset, or an interaction rollout depending on the task setting

This formulation provides a common interface for evaluating heterogeneous systems while allowing each modality to retain its own output structure and control semantics.

The task suite separates world generation into complementary regimes that probe different capabilities:

  • Static world generation — extends or transforms a scene under camera movement, viewpoint change, or scene-level layout specification while preserving geometry, appearance, and semantic consistency
  • Dynamic world generation — in-scene temporal evolution (object motion, human activity, environmental dynamics, state transitions), often under fixed or controlled camera settings to isolate dynamics from viewpoint changes
  • Interaction-conditioned generation — generated futures remain consistent with explicit actions, object controls, or intervention instructions

By decomposing tasks in this way, WorldAtlas distinguishes perceptual quality from spatial consistency, temporal dynamics, and controllability, enabling systematic comparison across video generators, 3D world generators, 4D dynamic scene models, and action-conditioned world models.

WorldAtlas treats each benchmark instance as a conditional next-world generation problem. A task is specified by the tuple (C,N,L)(C, N, L), and an evaluated system produces a generated world observation OO. The evaluation objective is to quantify how well OO satisfies the contract implied by (C,N,L)(C, N, L) across complementary capability dimensions—perceptual quality, spatial and temporal consistency, controllability, closed-loop memory, physical plausibility (including temporal calibration).

Formally, for sample ii and metric jj, the benchmark seeks normalized scores

s~i,j∈[0,1],\tilde{s}_{i,j} \in [0, 1],

together with suite-level and dimension-level aggregates over eligible outputs. No single benchmark-wide scalar is defined; comparability requires matched task regimes, sampling protocol, and prediction completeness.

Inputs.

SymbolRole
C={X,P}C = \{X, P\}Current evidence: visual observation XX (image or video) and optional textual description PP
NNNavigation or target-world specification: the intended future scene, dynamic change, or interaction outcome
LLLayout and control specification: camera trajectory, viewpoint change, navigation path, object-level control, action instruction, or fixed-camera constraint
OOGenerated world observation produced by the evaluated generative system

Auxiliary supervision—reference rollouts, pose trajectories, segmentation masks, or instruction labels—may accompany a sample when available. These sidecars gate which metrics are applicable but are not part of the core generative contract.

Outputs.

OutputDescription
Per-sample metric scoresRaw measurements and their normalized counterparts for each applicable metric
Eligibility statusOne official public tier; retain per-sample N/A, missing and failed states
Aggregated benchmark scoresSuite-level per-metric averages and dimension-level roll-ups computed only over eligible official scores
Historical summariesValidate and combine both legacy summary channels as official metrics

The benchmark applies a fixed five-stage procedure to every evaluated sample.

Stage 1 — Task instantiation. Each sample materializes (C,N,L)(C, N, L) under a suite-specific conditioning regime (static world generation, dynamic world generation, or interaction-conditioned generation). The conditioning contract determines which portion of the evidence serves as input and which temporal region of the rollout is evaluated.

Stage 2 — Protocol alignment. A deterministic temporal sampling protocol selects matched evaluation support on the prediction and, when required, on the reference stream. Anchor positions and motion windows ensure that prediction–reference comparisons occur on aligned time indices rather than arbitrary frame offsets.

Stage 3 — Metric computation. For each configured metric mjm_j, an evaluator computes a raw measurement

ri,j=fj ⁣(Oi, Ci, Ai),r_{i,j} = f_j\!\left(O_i,\, C_i,\, \mathcal{A}_i\right),

where Ai\mathcal{A}_i denotes the auxiliary supervision available for sample ii. Metrics span quality, spatial consistency, temporal dynamics, controllability, closed-loop memory, physics realism, and physical-time calibration. Some metrics operate on OO alone; others require paired reference content or annotated sidecars. Memory metrics are self-referential: the reference for a revisit frame is an earlier frame the model itself generated.

Stage 4 — Normalization and validity. Normalize raw measurements using each metric’s specified mapping. All documented and leaderboard metrics are official. Preserve N/A, failures and missing values with their reasons; do not automatically zero-fill them.

Stage 5 — Aggregation. Aggregate validated results with each metric’s specified denominator, including valid values stored under legacy diagnostic labels. Dimension and Overall scores follow the published combination rules; do not mix different protocols.

All metrics share one public status. Compare models using applicable inputs, valid measurements and verified coverage.