EMBER-Bench

Benchmarking Cross-Event Causal Memory
in Long-Horizon Embodied Tasks

* Equal contribution† Corresponding authors

an Egocentric Memory Benchmark for Embodied Reasoning

The past is
part of the present.

Which past event matters for the next action? EMBER-Bench tests whether models turn the continuing consequences of history into decisions.

An unfolding field of causal memory Conceptual schematic. Signals move between past events, lasting consequences, a current decision, and the next action. This is an illustration, not benchmark data. PAST EVENTLASTING CONSEQUENCE CURRENTDECISION NEXT ACTIONCHANGED STATETASK PROGRESS MEMORY FIELDt → t + 1
HISTORY → STATE → ACTIONCONCEPTUAL SCHEMATIC
189
Newly recorded household tasks
699
History-dependent questions
4
Memory scenarios, L1–L4
16
Evaluated multimodal models
01 / The question behind the benchmark

Memory is more
than a record.

Past events keep changing the world after they leave view. A completed prerequisite, an unresolved mistake, or a moved object can constrain what should happen next.

EMBER-Bench asks models to connect those events to the present: predict the next action, then trace that action back to its historical cause. The challenge is cross-event causal reasoning, beyond historical retrieval.

Explore the evaluation
A MEMORY DEPENDENCY, IN TWO DIRECTIONS
Earlier / Historical event

The mug was moved.

While the performer is elsewhere, another person places the mug inside a cupboard.

Now / Persistent state

The countertop is empty.

The task is to prepare tea. The current view alone does not reveal where the mug went.

Next / History-grounded action

Retrieve the mug.

Its changed location determines the next useful action: open the cupboard.

PREDICTION →

Given the goal and history, what should happen next?

Select the appropriate next action from four candidates. The earlier relocation makes retrieving the mug from the cupboard necessary.

ILLUSTRATIVE SCHEMATIC · This hypothetical example explains the task; it is not a released dataset item.

Original benchmark introduction showing task filmstrips for four memory scenarios and the paired next-action prediction and causal traceback questions.
Benchmark introduction. P asks what to do next; C identifies the historical interval supporting the given action.Open original PDF ↗
02 / Design

Four ways the past
shapes what comes next.

Newly designed and recorded egocentric household tasks, with fine-grained action and causal-chain annotations.

L1

Task Progress Memory

Track completed steps and prerequisites, even when task progress is no longer visible.

L2

Self-induced Anomaly Memory

Remember the lasting consequences of your own mistakes and the recovery required.

L3

External Intervention Memory

Update object location, identity, or availability after another person changes the scene.

L4

Compound Long-Horizon Memory

Maintain multiple historical constraints across intervening activities, anomalies, and interventions.

Original five-stage construction pipeline: task scenario design and capture, data filtering, annotation and QA construction, iterative quality control, and annotated output.
Construction pipeline: task scenario design and capture, filtering, annotation and QA construction, iterative quality control, and annotated output.Open original image ↗

P   Next-action prediction · 548 questions

Given a task goal, pre-decision video history, and current frame, select the next action from four candidates across L1–L4.

C   Causal traceback · 151 questions

Given the reference next action, select the relevant historical interval from four video-labeled candidates. Eligible L2–L4 questions are paired with prediction questions.

Fixed decision points in recorded human activities · Four-option multiple choice · Accuracy (%) · Offline reasoning evaluation

03 / What the results reveal

A visible gap.
An unresolved challenge.

See all 16 models
Video-history setting / Overall accuracy
Human reference · mean of 2 evaluators98.3%
Best evaluated model · Gemini 3.8 Flash61.2%
16-model average · Overall41.3%

Manuscript Table 2 · Accuracy across 699 questions. The 16-model average is the unweighted mean of their reported Overall accuracies. Human performance is a separate reference and is excluded from model ranking and the model average.

Remembering the cause does not ensure the right next action.

At matched decision points, correct traceback is not associated with correct next-action prediction in the reported analysis.

The content of memory matters.

Across six models, action logs add +1.6 points on average. Privileged cause-and-consequence annotations add a further +13.0 points on top of the logs.

The annotations diagnose information needs; they provide privileged causal information.

Compound memories are especially demanding.

12 of 16 models obtain their lowest next-action accuracy in L4, where multiple distant dependencies must be maintained together.

Original six-panel analysis figure showing history-length and event-distance breakdowns, diagnostic information ablations, and paired prediction and traceback outcomes.
History-length and event-distance analyses, diagnostic information ablations, and paired P/C outcomes. Annotation ablations use six models and 151 paired prediction questions with privileged causal information; these diagnostic results are separate from the 699-question main leaderboard. The original figure labels C as cause attribution (causal traceback).Open original PDF ↗
04 / Dataset coverage

Everyday tasks.
Distant dependencies.

189 newly recorded household tasks spanning scenes and activities, with diverse history lengths and causal dependencies.

Original dataset diversity figure showing household scenes and activities, input durations, action vocabulary, questions per memory level, and task lengths.
Dataset diversity across household scenes and activities, input durations, action vocabulary, memory levels, and task lengths.Open original PDF ↗

Human-led collection and annotation

The research team designed and recorded the tasks and annotated actions and causal relationships. Model-assisted QA verification and iterative human review check evidence, answer uniqueness, option bias, and input boundaries.

699 questions across four memory scenarios

L1: 205 · L2: 171 · L3: 188 · L4: 135. These are question counts, not task counts. The benchmark pairs 151 eligible prediction questions with causal traceback questions.

05 / Research resources

Study it. Evaluate it.
Build on it.

Explore the benchmark design and evaluation toolkit. The public dataset release is pending.

01 / BENCHMARK

Explore the benchmark

Four memory scenarios, paired reasoning directions, original task illustrations, and diagnostic analysis.

See the benchmark design
02 / EVALUATION

Run the benchmark

Evaluation code, input preparation, inference adapters, scoring, and reproducibility instructions.

Explore the evaluation guide
03 / DATASET

The next release

Full videos, questions, and annotations will be announced in the repository with access instructions.

Public dataset release pending