Community proposal aims to create a harness-only benchmark and leaderboard for LLM performance across real-world tasks
Read the original at old.reddit.com→There are a lot of LLM benchmarks but few, if any, harness benchmarks. I am thinking this would be a really good community project to build one. End goal: a leaderboard of harness performance (multiple axis) on a set...
Original headline: "Anyone interested in building a harness-only benchmark?"
Coverage timeline
- Aug 5, 10:54 UTC r/LocalLLaMA lead source Anyone interested in building a harness-only benchmark?