AR
arXiv CS.AI
7/22/2026

SAAG: Structured Agent Assessment and Grounding
Short summary
SAAG is a cascaded diagnostic framework that decomposes agent-calling evaluation into three sequential stages: registry conformance, structural completeness, and argument grounding. Each stage produces interpretable diagnostics that enable iterative self-repair without leaking ground-truth values. Evaluated on sub-4B-parameter models across registry sizes of 5-15 agents, structured feedback consistently improves argument precision and reduces value hallucination, though end-to-end F1 gains remain modest and model-dependent.
- •Three-stage evaluation: registry conformance, structural completeness, argument grounding
- •Stage-specific diagnostics enable targeted self-repair without ground-truth leakage
- •Structured feedback reduces value hallucination across sub-4B models, though F1 gains are modest
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
