About the tool
Framework and open registry of benchmarks for evaluating LLMs.
Evals is OpenAI's MIT-licensed framework for evaluating LLMs and LLM systems, shipping with a community registry of benchmark evals. It supports custom eval templates and model-graded evaluations that can be run against any model behind an API.
- Pricing
- Open source
- Category
- Evals
- Maker
- Not publicly supplied
- Provenance
- Indexed by Attest from public information
✓ Attested
✓ Attested