smevals - a small eval suite for evaluating models, prompts, and harnesses
A lightweight eval framework for practitioners who need fast, local benchmarking without the enterprise overhead.

Why it matters
smevals is a practical open tool for measuring model and prompt quality in production settings — it lowers the barrier for continuous evaluation, addressing a real gap between frontier benchmarks and day-to-day engineering.
The key facts
9 to knowOpen-source eval suite designed for models, prompts, and harnesses
Published by Simon Willison (Datasette creator, influential practitioner voice)
Targets practitioners doing local/rapid iteration, not just lab-scale benchmarking
Fills the gap between heavyweight benchmarks (SWE-bench, MMLU) and ad-hoc testing
smevals is a lightweight evaluation framework
Designed for evaluating models, prompts, and harnesses
Targets practitioners working with models iteratively
Published by Simon Willison (Datasette/LLM community authority)
Addresses the gap between full benchmarks and ad-hoc testing
Go to the source
Simon Willisonsimonwillison.net