Benchmark Overview
Task formulation
Section titled “Task formulation”Grounded in our operational definition of world models and the formal criterion introduced in Sec. 2, WorldAtlas formulates world generation as a set of conditional next-world generation tasks across video, 3D, and 4D modalities.
Task specification
Section titled “Task specification”Following a next-scene evaluation perspective, each task instance is specified by a tuple :
- — current evidence: an image or video observation and an optional textual description
- — target world specification: the next scene, dynamic change, object transformation, or interaction outcome to be generated
- — layout and control specification: camera trajectory, viewpoint change, navigation path, object-level control, action instruction, or fixed-camera constraint
Generative interface
Section titled “Generative interface”Given this specification, a model is required to generate an output artifact
where:
- denotes the evaluated generative system
- denotes model-specific preprocessing
- may be a video sequence, a 3D scene, a 4D dynamic asset, or an interaction rollout depending on the task setting
This formulation provides a common interface for evaluating heterogeneous systems while allowing each modality to retain its own output structure and control semantics.
Task regimes
Section titled “Task regimes”The task suite separates world generation into complementary regimes that probe different capabilities:
- Static world generation — extends or transforms a scene under camera movement, viewpoint change, or scene-level layout specification while preserving geometry, appearance, and semantic consistency
- Dynamic world generation — in-scene temporal evolution (object motion, human activity, environmental dynamics, state transitions), often under fixed or controlled camera settings to isolate dynamics from viewpoint changes
- Interaction-conditioned generation — generated futures remain consistent with explicit actions, object controls, or intervention instructions
By decomposing tasks in this way, WorldAtlas distinguishes perceptual quality from spatial consistency, temporal dynamics, and controllability, enabling systematic comparison across video generators, 3D world generators, 4D dynamic scene models, and action-conditioned world models.
Detailed Algorithm
Section titled “Detailed Algorithm”Problem formulation
Section titled “Problem formulation”WorldAtlas treats each benchmark instance as a conditional next-world generation problem. A task is specified by the tuple , and an evaluated system produces a generated world observation . The evaluation objective is to quantify how well satisfies the contract implied by across complementary capability dimensions—perceptual quality, spatial and temporal consistency, controllability, closed-loop memory, physical plausibility (including temporal calibration).
Formally, for sample and metric , the benchmark seeks normalized scores
together with suite-level and dimension-level aggregates over eligible outputs. No single benchmark-wide scalar is defined; comparability requires matched task regimes, sampling protocol, and prediction completeness.
Inputs and outputs
Section titled “Inputs and outputs”Inputs.
| Symbol | Role |
|---|---|
| Current evidence: visual observation (image or video) and optional textual description | |
| Navigation or target-world specification: the intended future scene, dynamic change, or interaction outcome | |
| Layout and control specification: camera trajectory, viewpoint change, navigation path, object-level control, action instruction, or fixed-camera constraint | |
| Generated world observation produced by the evaluated generative system |
Auxiliary supervision—reference rollouts, pose trajectories, segmentation masks, or instruction labels—may accompany a sample when available. These sidecars gate which metrics are applicable but are not part of the core generative contract.
Outputs.
| Output | Description |
|---|---|
| Per-sample metric scores | Raw measurements and their normalized counterparts for each applicable metric |
| Eligibility status | One official public tier; retain per-sample N/A, missing and failed states |
| Aggregated benchmark scores | Suite-level per-metric averages and dimension-level roll-ups computed only over eligible official scores |
| Historical summaries | Validate and combine both legacy summary channels as official metrics |
Evaluation pipeline
Section titled “Evaluation pipeline”The benchmark applies a fixed five-stage procedure to every evaluated sample.
Stage 1 — Task instantiation. Each sample materializes under a suite-specific conditioning regime (static world generation, dynamic world generation, or interaction-conditioned generation). The conditioning contract determines which portion of the evidence serves as input and which temporal region of the rollout is evaluated.
Stage 2 — Protocol alignment. A deterministic temporal sampling protocol selects matched evaluation support on the prediction and, when required, on the reference stream. Anchor positions and motion windows ensure that prediction–reference comparisons occur on aligned time indices rather than arbitrary frame offsets.
Stage 3 — Metric computation. For each configured metric , an evaluator computes a raw measurement
where denotes the auxiliary supervision available for sample . Metrics span quality, spatial consistency, temporal dynamics, controllability, closed-loop memory, physics realism, and physical-time calibration. Some metrics operate on alone; others require paired reference content or annotated sidecars. Memory metrics are self-referential: the reference for a revisit frame is an earlier frame the model itself generated.
Stage 4 — Normalization and validity. Normalize raw measurements using each metric’s specified mapping. All documented and leaderboard metrics are official. Preserve N/A, failures and missing values with their reasons; do not automatically zero-fill them.
Stage 5 — Aggregation. Aggregate validated results with each metric’s specified denominator, including valid values stored under legacy diagnostic labels. Dimension and Overall scores follow the published combination rules; do not mix different protocols.
Design principles
Section titled “Design principles”All metrics share one public status. Compare models using applicable inputs, valid measurements and verified coverage.