SAAG: Structured Agent Assessment and Grounding
Read the original at arxiv.org→arXiv:2607.18245v1 Announce Type: new Abstract: Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or...