AR
arXiv CS.AI
6/30/2026

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes
Short summary
GPTNT introduces a benchmark for multimodal agent collaboration using Keep Talking and Nobody Explodes, requiring real-time communication under time pressure and information asymmetry. Unlike existing turn-based evaluations, it measures state tracking, ambiguity handling, and error recovery in asynchronous settings. State-of-the-art models fail to defuse a single bomb, while humans succeed, revealing critical gaps in real-world collaboration.
- •Benchmark requires real-time agent coordination under time pressure, not turn-based exchanges
- •Tests state tracking, ambiguity handling, and error recovery in procedurally generated scenarios
- •All tested models fail where humans succeed, showing significant capability gaps
Generated with AI, which can make mistakes.
Is this a good recommendation for you?