Main results / Video-history setting · V

Memory, measured.

All results from manuscript Table 2. Accuracy (%) on next-action prediction and causal traceback, with every reported scenario breakdown.

Download CSV
16 models699 questions61.2% best model overall98.3% human overallVideo-history setting · Paper Table 2
16 of 16 models
EMBER-Bench video-history accuracy. Human reference is unranked. Ten score columns: overall, five prediction columns, and four traceback columns.
#ModelNext-action prediction · PCausal traceback · C
—Human (mean of 2)Human reference 98.397.8100.096.997.794.6100.0100.0100.0100.0
1Gemini 3.8 FlashClosed-source 61.258.468.848.757.849.071.560.388.360.6
2Gemini 3.1 Pro PreviewClosed-source 51.550.761.047.840.646.154.348.368.339.4
3Doubao-Seed-2.1-ProClosed-source 49.845.455.638.943.834.365.655.276.763.6
4Qwen3.8 MaxClosed-source 47.544.252.239.844.532.459.656.968.348.5
5GPT-6 SolClosed-source 46.244.257.141.632.835.353.651.758.348.5
6Qwen3.8 Omni FlashClosed-source 45.644.355.136.343.832.450.343.156.751.5
7Qwen3.5 397B-A17BOpen-source 45.443.852.742.538.334.351.048.361.736.4
8Qwen3.5 35B-A3BOpen-source 39.538.545.934.535.232.443.041.450.033.3
9Qwen3.5 122B-A10BOpen-source 38.235.845.429.234.425.547.044.856.733.3
10Qwen3.8 27BOpen-source 38.136.742.431.038.329.443.041.448.336.4
11Qwen3.8 FlashClosed-source 36.834.739.031.935.927.544.443.150.036.4
12MiMo V2.6 ProOpen-source 35.833.439.530.133.624.544.441.451.736.4
13Qwen3-VL 32BOpen-source 34.532.541.027.428.925.541.743.146.730.3
14MiMo V2.6 FlashOpen-source 32.328.532.227.427.323.546.444.851.739.4
15Qwen3-VL 8BOpen-source 29.029.933.730.131.220.625.829.328.315.2
16Qwen3-VL 235B-A22BOpen-source 28.926.632.727.418.823.537.137.940.030.3

Amber values mark the best model score in each column. P: 548 questions across L1–L4. C: 151 questions across L2–L4. Unanswered items after retries are incorrect; random-choice accuracy is 25%. Human reference is the mean of two evaluators.

Ranks: Overall across all 16 models. Ties share rank. Human reference remains unranked.

Evaluate a model. Share a result.

The evaluation guide covers the data contract, inference, and scoring. Public dataset access is pending. New results require evidence and review before inclusion.

Result review guide
Reading the results

Two directions.
Four kinds of memory.

Evaluation protocol

What the metrics measure

Prediction (P) selects a next action from four candidates given the goal, pre-decision video history, and current frame.

Traceback (C) receives the reference next action and selects the historical interval that contains its supporting event. Four candidate intervals are labeled on the video.

Overall pools correctness across all 699 questions. L1 contains prediction questions only. C is paired with a subset of P in L2–L4.

The scenario breakdown

L1 / TASK PROGRESSCompleted steps and prerequisites.
205 P · 0 C
L2 / SELF-INDUCED ANOMALYLasting mistakes and recovery obligations.
113 P · 58 C
L3 / EXTERNAL INTERVENTIONThird-party changes to objects or state.
128 P · 60 C
L4 / COMPOUND LONG-HORIZONMultiple distant causal constraints.
102 P · 33 C