UserToolBench tests personalized decision making in tool-use LLMs by inferring latent user preferences from interaction history
Read the original at arxiv.org→arXiv:2608.10042v1 Announce Type: new Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or...
Original headline: "UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs"
Coverage timeline
- Aug 12, 04:00 UTC arXiv cs.LG lead source UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs