HomePeopleCompaniesAI ModelsOpen SourceAgentsResearchApps
AllModelsInferenceToolingImage and Video

llm-bench-runner: Run and Compare LLM Benchmarks Locally Without Cloud

The brief

llm-bench-runner is an open-source tool for running MMLU, HumanEval, MATH, and GPQA benchmarks locally against any LLM, with visualization of results.

Key points

  1. It eliminates dependency on cloud benchmark services for practitioners who want reproducible local evaluation.
Read the original

Sources