Terminal-Bench-LILT proposes a multilingual coding benchmark with 300 tasks across ten languages to evaluate non-English software development challenges.
Read the original at arxiv.org→arXiv:2608.28641v1 Announce Type: new Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a...
Original headline: "Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture"
Coverage timeline
- Sep 1, 04:00 UTC arXiv cs.CL lead source Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture