About the tool
EleutherAI's unified framework for benchmarking generative language models.
The lm-evaluation-harness provides a unified interface for testing language models across hundreds of standardized evaluation tasks, and is the backend used by the Open LLM Leaderboard. It supports config-based task creation and a wide range of model backends under an MIT license.
- Pricing
- Open source
- Category
- Evals
- Maker
- Not publicly supplied
- Provenance
- Indexed by Attest from public information
✓ Attested
✓ Attested