Preprint · 2026Long-term egocentric memory

EgoMonthA Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

What can a model remember after a month of everyday life?
Test how it retains, retrieves, and connects experiences across days.

Weitao Chen1,*Jiaxin Hu1,*Tianyidan Xie1Yang Li1Yuyi Qian1Banghao Xu1Ziheng Tang1Shenyi Wang1Mingyue Yu1Duo Li1Jiacheng Shi1Gao Wang1Zhan Xu2Zhicheng Qiu2Xuanfu Li2Jian Yang1Lanjun Wang3,✉Zili Yi1,✉
1Nanjing University2Huawei Technologies Co., Ltd.3Tianjin University

* Equal contribution · ✉ Corresponding authors: Lanjun Wang, Zili Yi

Original EgoMonth overview: a month-long timeline of first-person recordings alongside everyday scenes including dining, commuting, working, and shopping.
Daily recordings share places, people, and routines. EgoMonth asks models to connect this experience across time.Explore Figure 1 ↗
301 hEveryday first-person video738 clips · 20 participants
20–120Days per participantLongitudinal recording span
1,443Human-crafted questionsFour-option multiple-choice QA
14Tasks, three cognitive levelsPatterns → episodes → reasoning

The research question

Daily life.
Lasting memory?

See the questions ↓

Remembering a month of daily life means keeping track of recurring places, sparse events, and subtle changes. A useful memory must distinguish what usually happens from what happened at a particular time, then connect observations when a question requires more than one episode.

EgoMonth brings this challenge to first-person video understanding. Its recordings follow independent daily routines across 20–120 days. The benchmark evaluates both questions grounded in a single recording and questions that combine evidence across multiple videos.

22.4 ppThe reported gap between Gemini 2.5 Pro and the human reference on macro-average accuracy: 71.8% vs. 94.2%.
View paper results · Explore the comparison ↓

Inside the benchmark

Experience with continuity.

The same participant returns to familiar settings over days and weeks, with important evidence scattered among ordinary routines.

01 / COLLECT

Record daily life

Smartphones and action cameras capture indoor and outdoor activities. The retained corpus spans 20 participants and 738 clips.

20–120 days per participant

02 / CURATE

Keep usable context

Screen recordings for quality, viewpoint, and continuity. Anonymize sensitive regions and check the results manually.

18,072 minutes of retained video

03 / ANNOTATE

Write evidence-based QA

Annotators inspect events, object states, and temporal dependencies to write questions with four plausible answer options.

Fully human-written questions

04 / REVIEW

Verify the evidence

Groups of three reviewers check answer correctness and distractors, revisiting the recordings when they disagree.

1,443 QA pairs across 14 tasks

Recording span describes the interval across days, not uninterrupted 24-hour footage. Average clip length is approximately 24.5 minutes. Dataset construction and statistics.

Explore the recording and activity distributions
Figure 3(a) · Recording distributionEnlarge ↗
Figure 3(a): distribution of recording span, clip duration, and indoor versus outdoor scenes.
The paper summarizes the recording span, clip duration, and scene environment in a nested distribution plot.
Figure 4 · Everyday activitiesEnlarge ↗
Figure 4: activity categories include study and work, food and dining, shopping, leisure, household chores, and transportation.
Six activity groups cover work and study, transport, dining, shopping, leisure, and household chores.

Questions in context

Find the evidence.
Connect the experience.

Two original examples from the paper illustrate the difference between following one activity and finding a pattern across recordings.

Procedure planning · Level 3

How did the wearer prepare the ingredients?

The answer depends on the order of several actions. Retrieve the relevant moments in a cooking recording and assemble them into the observed procedure.

Original single-video QA example with a cooking question, four answer options, and timestamped evidence showing vegetable selection, washing, ingredient retrieval, stir-frying, and serving.
The reference answer follows five steps: select vegetables, wash them, take ingredients from the refrigerator, stir-fry, and serve. The evidence is distributed within one recording.

Habit inference · Level 1

Which leisure habit recurs across days?

A single clip cannot establish a persistent routine. Combine observations from multiple recordings to identify the behavior that repeats across different settings.

Original cross-video QA example linking recordings over two weeks. Repeated observations of phone use in several locations support the reference answer.
The reference answer identifies recurring phone use across locations such as a bed, chair, and sofa. Evidence spans roughly two weeks of recordings.

14 tasks / 3 levels

From patterns
to precise recall.

A progressive hierarchy separates stable behavioral patterns, retrieval of specific episodes, and reasoning that depends on multiple pieces of evidence.

Schema Consolidation2 tasks · 154 questions

Repeated experiences provide multiple chances to recover a stable pattern. These tasks ask whether a model can infer regularities across weeks of daily life.

Habit Inference

Identify routines supported by repeated observations across weeks.

138 questions · 9.6% of benchmark

Personality Inference

Infer enduring behavioral traits from the recorded evidence.

16 questions · 1.1% of benchmark

Episodic Indexing6 tasks · 873 questions

Find a specific detail, place, object state, or point in time. Similar-looking episodes make precise retrieval harder than remembering the general story.

Detail Retrieval

Recall a fine-grained visual detail from a particular episode.

287 questions · 19.9% of benchmark

Spatial Relation

Recover how objects are positioned relative to one another.

51 questions · 3.5% of benchmark

Self-localization

Identify the wearer’s location at a specified moment.

85 questions · 5.9% of benchmark

Temporal Ordering

Arrange observed events in their correct chronological order.

127 questions · 8.8% of benchmark

Event Time

Locate when a particular event happened in the recordings.

117 questions · 8.1% of benchmark

Object Location

Retrieve where an object was last observed.

206 questions · 14.3% of benchmark

Cascading Reasoning6 tasks · 416 questions

Combine multiple retrieved observations into a count, route, procedure, or spatial relation. A missed observation or incorrect intermediate state can change the final answer.

Procedure Planning

Connect the steps of a multi-stage activity in sequence.

126 questions · 8.7% of benchmark

Event Counting

Accumulate occurrences of an event across the evidence.

125 questions · 8.7% of benchmark

Object Counting

Count object instances over the full video context.

57 questions · 4.0% of benchmark

Route Reasoning

Reconstruct a route from successive places and movements.

42 questions · 2.9% of benchmark

Cross-view Spatial Reasoning

Relate spatial evidence across viewpoints and recording days.

30 questions · 2.1% of benchmark

Direction Judgement

Determine orientation from the wearer’s changing viewpoint.

36 questions · 2.5% of benchmark
Task definitions: Table 1. Question counts: Table 5.Question counts · CSV ↓
View the full task distribution
Figure 3(b) · QA distribution across cognitive levelsEnlarge ↗
Figure 3(b): QA distribution across the 14 tasks and three levels, with episodic indexing containing the largest share of questions.
Question frequencies are deliberately heterogeneous. The results report both a macro-average across tasks and accuracy across all questions.

Reported results

A persistent
memory gap.

Twelve evaluated models, fourteen tasks, and a human reference. Explore overall performance or compare a specific memory skill.

11 open-source models1 closed-source model4 answer optionsDirect answer matching

Paper results · Accuracy (%)
Higher is better

Macro-average accuracy

  1. HumanReference94.2%
  2. Gemini 2.5 ProClosed71.8%
  3. Qwen2.5-VL-32BOpen58.0%
  4. MiniCPM-V 4.5Open56.0%
  5. Qwen2-VLOpen54.5%
  6. Qwen3-VL-30B-A3BOpen53.0%
  7. Qwen3-VL-8BOpen51.4%
  8. VITA-1.5Open51.3%
  9. VideoLLaMA3Open50.3%
  10. LLaVA-NeXT-VideoOpen40.6%
  11. Chat-UniVi-V1.5Open39.5%
  12. ST-LLMOpen38.8%
  13. ShareGPT4VideoOpen37.1%
Dashed line: 25% random chanceHuman: mean of three annotators

Gemini 2.5 Pro: 71.8% · Human reference: 94.2% · Gap: 22.4 percentage points.

Every task, every evaluated model

All values below reproduce the paper’s Table 3, including its reported aggregate scores and model input configurations.

Scroll sideways to compare all models →

Table 3 · arXiv v1. Accuracy (%), higher is better. Human values are a reference. A dash means the paper does not specify the value.
Task / metricChat-UniVi-V1.5LLaVA-NeXT-VideoMiniCPM-V 4.5Qwen2-VLQwen2.5-VL-32BQwen3-VL-8BQwen3-VL-30B-A3BShareGPT4VideoST-LLMVideoLLaMA3VITA-1.5Gemini 2.5 ProHuman
Input frames2566425625625625625664256512161fps
Parameters7B7B8B7B32B8B30B8B7B7B
Level 1 · Schema Consolidation
Habit Inference58.772.582.675.481.975.481.950.065.269.678.384.897.8
Personality Inference50.062.568.856.262.568.837.550.056.250.056.281.393.8
Level 2 · Episodic Indexing
Detail Retrieval34.138.366.558.961.754.456.834.834.855.451.677.795.1
Spatial Relation56.949.058.854.956.951.051.051.049.049.054.986.394.1
Self-localization29.441.254.152.949.452.952.931.835.243.548.282.497.6
Temporal Ordering40.945.770.159.864.655.962.236.240.162.257.575.695.3
Event Time44.446.157.342.747.941.950.445.342.741.949.660.795.7
Object Location51.950.561.661.664.155.357.344.758.254.951.567.096.1
Level 3 · Cascading Reasoning
Procedure Planning46.029.468.268.278.671.470.630.246.060.366.781.794.4
Event Counting8.816.840.036.843.238.434.416.88.040.841.652.094.4
Object Counting33.329.838.652.650.931.647.429.831.649.140.473.794.7
Route Reasoning35.723.835.752.452.428.645.230.928.645.240.564.390.5
Cross-view Spatial Reasoning23.330.043.340.050.046.746.726.720.040.033.360.093.3
Direction Judgement38.933.338.950.047.247.247.241.727.841.747.258.386.1
Avg · macro39.540.656.054.558.051.453.037.138.850.351.371.894.2
Acc · micro39.941.760.657.060.853.756.736.940.853.153.672.695.1

Source: Table 3, arXiv v1. Models use different input budgets, following their evaluation configurations; this is not a controlled comparison at a shared frame budget. Avg is the reported macro-average across tasks; Acc is accuracy across all questions.

How answers are scored

Outputs are compared directly with the reference answers; no LLM judge is used. Four-option questions have a 25% chance baseline. The paper evaluates each model in a single pass without ensembling.

How the human reference is measured

Three trained annotators, independent of QA creation, answer the full benchmark with unrestricted viewing and replay. Their reported average is 94.2% macro and 95.1% micro, with Fleiss’ κ = 0.78.

What the benchmark reveals

Recognition is only
the beginning.

The reported results point to weaknesses in keeping time, counting events, and joining spatial evidence over long contexts.

35.0 pp

Locate the right moment

On Event Time, Gemini 2.5 Pro scores 60.7%, compared with 95.7% for humans. Remembering the event’s general meaning does not guarantee a correct temporal index.

42.4 pp

Keep an accurate count

Event Counting reaches 52.0% for Gemini 2.5 Pro versus 94.4% for humans. Multiple open-source models fall below the 25% chance baseline.

33.3 pp

Connect different viewpoints

Cross-view Spatial Reasoning reaches 60.0% for Gemini 2.5 Pro versus 93.3% for humans. Stable spatial relationships remain difficult across changing observations.

Gaps above are percentage-point differences calculated from Table 3. The interpretation follows the paper’s analysis.

Scope and interpretation
  • The current benchmark includes 20 participants. It does not establish performance across every demographic or environment.
  • Coverage extends to 120 days per participant; seasonal and year-long memory are future directions.
  • The human reference has unrestricted access to the full recordings, while models use the input configurations reported in the paper.
  • Frame-budget and model-size comparisons are observational. They do not isolate the causal effect of adding frames or parameters.

Use EgoMonth

Research access
and resources.

The paper describes a research-only access policy for the benchmark. Metadata and QA pairs can be shared for research; raw videos require additional approval. Contact the authors about availability and access. Redistribution and commercial use are prohibited under the policy described in the paper.

Reference

Cite EgoMonth.

arXiv:2608.13113 · 2026 preprint. The linked PDF includes the full paper and appendices.

@misc{chen2026egomonth,
  title = {EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory},
  author = {Chen, Weitao and Hu, Jiaxin and Xie, Tianyidan and Li, Yang and Qian, Yuyi and Xu, Banghao and Tang, Ziheng and Wang, Shenyi and Yu, Mingyue and Li, Duo and Shi, Jiacheng and Wang, Gao and Xu, Zhan and Qiu, Zhicheng and Li, Xuanfu and Yang, Jian and Wang, Lanjun and Yi, Zili},
  year = {2026},
  eprint = {2608.13113},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  doi = {10.48550/arXiv.2608.13113},
  url = {https://arxiv.org/abs/2608.13113}
}

Paper figure