FindStatBench: evaluating large language models on combinatorial code synthesis
Read the original at arxiv.org→arXiv:2607.18260v1 Announce Type: new Abstract: We introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis. Built from FindStat, it contains 2,329 tasks...
Original headline: "FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis"