lm-eval-harness
The evaluation harness that powers the Hugging Face Open LLM Leaderboard, supporting 60+ standardized benchmarks. If you are fine-tuning or comparing language models, this is the tool the community trusts for apples-to-apples evaluation.
AI ToolFree Tier: MITCompany: EleutherAICategory: Safety & EthicsOpen Source: YesQuick Start: Install with pip install lm-eval → Run lm_eval --model hf --model_args pretrained=your-model --tasks hellaswag → Compare benchmark scores across modelsVisit lm-eval-harness