lm-eval-harness

The evaluation harness that powers the Hugging Face Open LLM Leaderboard, supporting 60+ standardized benchmarks. If you are fine-tuning or comparing language models, this is the tool the community trusts for apples-to-apples evaluation.

AI ToolFree Tier: MITCompany: EleutherAICategory: Safety & EthicsOpen Source: YesQuick Start: Install with pip install lm-eval → Run lm_eval --model hf --model_args pretrained=your-model --tasks hellaswag → Compare benchmark scores across models

lm-eval-harness

60+ benchmarks — powers the Open LLM Leaderboard

Visit lm-eval-harness

The evaluation harness that powers the Hugging Face Open LLM Leaderboard, supporting 60+ standardized benchmarks. If you are fine-tuning or comparing language models, this is the tool the community trusts for apples-to-apples evaluation.

FREE TIER
MIT
COMPANY
EleutherAI
CATEGORY
Safety & EthicsEvaluation
OPEN SOURCE
Yes
PRIVACY
Standard
TAGS
evaluationbenchmarks

QUICK START

Install with pip install lm-eval → Run lm_eval --model hf --model_args pretrained=your-model --tasks hellaswag → Compare benchmark scores across models

BEST FOR

  • LLM researchers needing standardized, reproducible benchmarks for model comparison.
  • Developers fine-tuning LLMs who require objective metrics for performance tracking.
  • Teams contributing to or leveraging the Hugging Face Open LLM Leaderboard.

NOT FOR

  • Non-technical users seeking quick, qualitative feedback on their AI models.
  • Anyone not working with large language models or requiring application-specific metrics.
Finding similar tools…
← Back to all tools