Dev.to
7/30/2026

OpenAI Says AI Benchmark Scores Depend on Harnesses, Budgets, and Memory Design
Short summary
OpenAI published guidance urging evaluators to treat benchmark scores as measurements of a model within a specific setup, not as fixed capability ceilings. The harness configuration, compute budget, available tools, and memory strategy all materially affect observed performance, especially on long-running multi-step agent tasks. OpenAI cites GPT-5.5 cyber-range evaluations where context compaction improved results as a concrete example of why these variables must be disclosed.
- •OpenAI says benchmark scores reflect a model in a particular harness, not standalone capability
- •Harness, budget, tools, and memory design are core variables evaluators must disclose
- •GPT-5.5 cyber-range evaluations showed performance gains from context compaction across turns
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



