AR
arXiv CS.AI
7/24/2026

Benchmarking the Personalization Capabilities of Large Language Models
Short summary
This paper benchmarks LLM personalization capabilities through a Bayesian Persuasion framework applied to sales outreach, releasing SDR-Bench with 6,279 customer success stories across 22 industries. Frontier LLMs and deep-research agents show a consistent personalization plateau, with no model statistically separating successful from unsuccessful outreach on a Fortune 100 cohort. A field deployment with 12 sales reps validated the framework, with 48% of model-generated content rated immediately useful and senior-expert agreement at Pearson 0.82.
- •SDR-Bench: 6,279 customer success stories for benchmarking LLM personalization in sales
- •Frontier LLMs show a personalization plateau — none distinguish successful from unsuccessful outreach
- •Field validation with 12 sales reps: 48% of AI content rated immediately useful, expert agreement at 0.82
Generated with AI, which can make mistakes.
Is this a good recommendation for you?