Dev.to
7/30/2026

OpenAI Says Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score
Short summary
OpenAI reports that two Responses API settings—retained reasoning and compaction—raised GPT-5.6 Sol's ARC-AGI-3 public-set score from 13.3% to 38.3% while reducing output tokens roughly sixfold. The improvement demonstrates that agent harness configuration, not just model weights, materially affects benchmark outcomes and operational cost. Developers building multi-step agents should test the full application stack, not just model selection, and use the Responses API with both settings enabled.
- •Retained reasoning and compaction settings tripled ARC-AGI-3 score from 13.3% to 38.3%
- •Output tokens reduced ~6x with the new configuration
- •Benchmark scores reflect harness configuration, not just model capability—test the full stack
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


