Benchy proposes a universal language and execution engine for task-oriented AI benchmarks, using YAML-encoded benchmarks defined by programs, scoring functions, and datasets.
Read the original at arxiv.org→arXiv:2609.30550v1 Announce Type: new Abstract: Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset,...
Original headline: "Benchy: towards a universal language for task-oriented AI benchmarks"
Coverage timeline
- Sep 28, 04:00 UTC arXiv cs.AI lead source Benchy: towards a universal language for task-oriented AI benchmarks