PB-GRPO: Learning socially adaptive LLM agents from persona-driven simulation with preference-batched GRPO
Read the original at arxiv.org→arXiv:2610.04132v1 Announce Type: new Abstract: Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing...
Original headline: "PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO"
Coverage timeline
- Oct 6, 04:00 UTC arXiv cs.CL lead source PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO