Smevals: a small eval suite for evaluating models, prompts, and harnesses
Read the original at simonwillison.net→smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about...
Original headline: "smevals - a small eval suite for evaluating models, prompts, and harnesses"
Coverage timeline
- Jul 31, 21:15 UTC Simon Willison lead source smevals - a small eval suite for evaluating models, prompts, and harnesses