smevals - a small eval suite for evaluating models, prompts, and harnesses
Simon Willison announced smevals, a new tool for running small eval suites across different model configurations and grading results. Developed with Jesse Vincent's Prime Radiant lab, it uses YAML-based evals, supports multiple models, and offers commands to run, grade, serve, and build static reports. Willison describes it as his third iteration on evals and feels it's right.
Provides developers a new tool to systematically evaluate models, prompts, and harnesses with reusable YAML-based evals.


