KC-Bench: a dynamic interactive benchmark for evaluating knowledge conflicts in LLM agents
Read the original at arxiv.org→arXiv:2609.03588v1 Announce Type: new Abstract: As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We...
Original headline: "KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.AI lead source KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
- Sep 4, 04:00 UTC arXiv cs.CL Large Language Models in Resolving Contextual Knowledge Conflicts