arXiv cs.CL
8/5/2026

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
Short summary
MemArena is a new ego-centric benchmark for evaluating on-device personal memory assistants, simulating 50 agents over 15 days with 10.3M dialog tokens. It tests six recall, reasoning, and trustworthiness dimensions across five open-weight readers and multiple memory backends. Key findings show memory-backend choice impacts accuracy more than reader scaling, permission-aware access fails universally, and search latency is manageable on edge devices.
- •MemArena benchmarks on-device personal memory assistants across 50 simulated agents over 15 days
- •Memory-backend choice matters more than reader scaling for content accuracy (+32.5pp gains observed)
- •Permission-aware access control fails across all tested backends, with Oracle leaking and others too timid
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

