smevals - a small eval suite for evaluating models, prompts, and harnesses
We've been building a new system for running evals against different models, prompts, and harnesses, with the goal of being able to identify the most appropriate small and inexpensive models for different categories of task.
