Spurious tool use: RL agents learn shortcut tool-selection policies based on superficial prompt cues rather than genuine task reasoning
Read the original at arxiv.org→arXiv:2609.16268v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly interleave natural language reasoning with external tools such as web search and code execution. These tool-use policies...
Original headline: "Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act"
Coverage timeline
- Sep 16, 04:00 UTC arXiv cs.CL lead source Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act