Meta and academic researchers release a benchmark for language models to find and fix bugs in large codebases before users encounter them.
Read the original at www.reddit.com→Hi! This is a new benchmark that I created together with other researchers at Meta, Stanford, Harvard, UW. Most benchmarks these days seem to test models to just fix a bug that I as a user already encountered. But...
Original headline: "New benchmark on LMs fixing bugs before users run into them"
Coverage timeline
- Oct 2, 16:05 UTC r/LocalLLaMA lead source New benchmark on LMs fixing bugs before users run into them