Back to feed
arXiv cs.CL
arXiv cs.CL
8/5/2026
MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

Short summary

MemArena is a new ego-centric benchmark for evaluating on-device personal memory assistants, simulating 50 agents over 15 days with 10.3M dialog tokens. It tests six recall, reasoning, and trustworthiness dimensions across five open-weight readers and multiple memory backends. Key findings show memory-backend choice impacts accuracy more than reader scaling, permission-aware access fails universally, and search latency is manageable on edge devices.

  • MemArena benchmarks on-device personal memory assistants across 50 simulated agents over 15 days
  • Memory-backend choice matters more than reader scaling for content accuracy (+32.5pp gains observed)
  • Permission-aware access control fails across all tested backends, with Oracle leaking and others too timid

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more