02 / PAPER-REPORTED PROTOCOL
Temporal evidence
and spatial agreement.
For both tasks, the paper reports success rate: a
retrieved frame must lie within the object's ground-truth temporal
span and have bounding-box IoU ≥ 0.3. The paper
also reports stricter IoU ≥ 0.5 results and per-horizon analysis.
-
STR annotations: tracking-assisted object
instances and state histories are manually verified. Annotators
identify valid temporal spans and representative-frame boxes;
candidates with failed tracking are skipped.
-
LOR annotations: annotators inspect
chronologically ordered observations in reverse to identify the
latest occurrence within each lookback window.
-
Interpretation: the tracking-assisted
construction and skipped instances are part of the benchmark's
scope. The published protocol should not be confused with a
released scoring program.
Where the observations come from
SMB is constructed from
EgoLife, whose source collection contains 300 hours from six participants
over seven days in a shared home. Individual sessions reach 50
hours. This is the source collection's scale, not the length of
every SMB query; LOR looks back at most 24 hours.
How SMB differs from Ego4D
The LTE paper also evaluates
Ego4D Natural Language Queries (NLQ) and
Visual Queries 2D (VQ2D). Those are established
Ego4D tasks, not part of SMB. SMB introduces STR and LOR using
EgoLife recordings to study long-term embodied agent memory, object
state history, and natural-language video retrieval.
Construction and annotation details: §4.1 and Appendix A of
arXiv:2609.04802v1. SMB does not evaluate general lifelong learning or robot
navigation.