llm-bench-runner: Run and Compare LLM Benchmarks Locally Without Cloud
The brief
llm-bench-runner is an open-source tool for running MMLU, HumanEval, MATH, and GPQA benchmarks locally against any LLM, with visualization of results.
Key points
- It eliminates dependency on cloud benchmark services for practitioners who want reproducible local evaluation.
Sources
- ingestgithub.com