Few-shot prompting shows task-dependent effects across 12 models and two tasks, with large gains on news classification but modest or negative changes on legal text.
Read the original at arxiv.org→arXiv:2609.15990v1 Announce Type: new Abstract: Few-shot prompting sometimes degrades language models instead of helping them, but why this happens is unknown. We evaluate 12 open-weight models on two Ukrainian...
Original headline: "Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures"
Coverage timeline
- Sep 16, 04:00 UTC arXiv cs.CL lead source Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures