The imperfective paradox is not necessarily in large language models; a benchmark failure before a model failure suggests conceptual and evaluation issues in reexamining the task.
Read the original at arxiv.org→arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer...
Original headline: "The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure"
Coverage timeline
- Aug 27, 04:00 UTC arXiv cs.CL lead source The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure