Back to feed
AR
arXiv CS.AI
6/30/2026
GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

Short summary

GPTNT introduces a benchmark for multimodal agent collaboration using Keep Talking and Nobody Explodes, requiring real-time communication under time pressure and information asymmetry. Unlike existing turn-based evaluations, it measures state tracking, ambiguity handling, and error recovery in asynchronous settings. State-of-the-art models fail to defuse a single bomb, while humans succeed, revealing critical gaps in real-world collaboration.

  • Benchmark requires real-time agent coordination under time pressure, not turn-based exchanges
  • Tests state tracking, ambiguity handling, and error recovery in procedurally generated scenarios
  • All tested models fail where humans succeed, showing significant capability gaps

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more