JOR-Bench: Japanese operations research benchmarks for large language models
Read the original at arxiv.org→arXiv:2607.16777v1 Announce Type: new Abstract: We present JOR-Bench, a collection of five Japanese-language benchmarks for evaluating the ability of large language models (LLMs) to formulate and solve operations...
Original headline: "JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models"