Memory, measured.
All results from manuscript Table 2. Accuracy (%) on next-action prediction and causal traceback, with every reported scenario breakdown.
| # | Model | Next-action prediction · P | Causal traceback · C | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| — | Human (mean of 2)Human reference | 98.3 | 97.8 | 100.0 | 96.9 | 97.7 | 94.6 | 100.0 | 100.0 | 100.0 | 100.0 |
| 1 | Gemini 3.8 FlashClosed-source | 61.2 | 58.4 | 68.8 | 48.7 | 57.8 | 49.0 | 71.5 | 60.3 | 88.3 | 60.6 |
| 2 | Gemini 3.1 Pro PreviewClosed-source | 51.5 | 50.7 | 61.0 | 47.8 | 40.6 | 46.1 | 54.3 | 48.3 | 68.3 | 39.4 |
| 3 | Doubao-Seed-2.1-ProClosed-source | 49.8 | 45.4 | 55.6 | 38.9 | 43.8 | 34.3 | 65.6 | 55.2 | 76.7 | 63.6 |
| 4 | Qwen3.8 MaxClosed-source | 47.5 | 44.2 | 52.2 | 39.8 | 44.5 | 32.4 | 59.6 | 56.9 | 68.3 | 48.5 |
| 5 | GPT-6 SolClosed-source | 46.2 | 44.2 | 57.1 | 41.6 | 32.8 | 35.3 | 53.6 | 51.7 | 58.3 | 48.5 |
| 6 | Qwen3.8 Omni FlashClosed-source | 45.6 | 44.3 | 55.1 | 36.3 | 43.8 | 32.4 | 50.3 | 43.1 | 56.7 | 51.5 |
| 7 | Qwen3.5 397B-A17BOpen-source | 45.4 | 43.8 | 52.7 | 42.5 | 38.3 | 34.3 | 51.0 | 48.3 | 61.7 | 36.4 |
| 8 | Qwen3.5 35B-A3BOpen-source | 39.5 | 38.5 | 45.9 | 34.5 | 35.2 | 32.4 | 43.0 | 41.4 | 50.0 | 33.3 |
| 9 | Qwen3.5 122B-A10BOpen-source | 38.2 | 35.8 | 45.4 | 29.2 | 34.4 | 25.5 | 47.0 | 44.8 | 56.7 | 33.3 |
| 10 | Qwen3.8 27BOpen-source | 38.1 | 36.7 | 42.4 | 31.0 | 38.3 | 29.4 | 43.0 | 41.4 | 48.3 | 36.4 |
| 11 | Qwen3.8 FlashClosed-source | 36.8 | 34.7 | 39.0 | 31.9 | 35.9 | 27.5 | 44.4 | 43.1 | 50.0 | 36.4 |
| 12 | MiMo V2.6 ProOpen-source | 35.8 | 33.4 | 39.5 | 30.1 | 33.6 | 24.5 | 44.4 | 41.4 | 51.7 | 36.4 |
| 13 | Qwen3-VL 32BOpen-source | 34.5 | 32.5 | 41.0 | 27.4 | 28.9 | 25.5 | 41.7 | 43.1 | 46.7 | 30.3 |
| 14 | MiMo V2.6 FlashOpen-source | 32.3 | 28.5 | 32.2 | 27.4 | 27.3 | 23.5 | 46.4 | 44.8 | 51.7 | 39.4 |
| 15 | Qwen3-VL 8BOpen-source | 29.0 | 29.9 | 33.7 | 30.1 | 31.2 | 20.6 | 25.8 | 29.3 | 28.3 | 15.2 |
| 16 | Qwen3-VL 235B-A22BOpen-source | 28.9 | 26.6 | 32.7 | 27.4 | 18.8 | 23.5 | 37.1 | 37.9 | 40.0 | 30.3 |
Amber values mark the best model score in each column. P: 548 questions across L1–L4. C: 151 questions across L2–L4. Unanswered items after retries are incorrect; random-choice accuracy is 25%. Human reference is the mean of two evaluators.
Ranks: Overall across all 16 models. Ties share rank. Human reference remains unranked.
Two directions.
Four kinds of memory.
What the metrics measure
Prediction (P) selects a next action from four candidates given the goal, pre-decision video history, and current frame.
Traceback (C) receives the reference next action and selects the historical interval that contains its supporting event. Four candidate intervals are labeled on the video.
Overall pools correctness across all 699 questions. L1 contains prediction questions only. C is paired with a subset of P in L2–L4.
The scenario breakdown
205 P · 0 C
113 P · 58 C
128 P · 60 C
102 P · 33 C