GitHub - david-g-3654/homebench: Benchmark your local LLMs: speed, memory, and quality, in one command. TUI leaderboard for Ollama, LM Studio, llama.cpp, and vLLM.
Pangram verdict · v3.3
We believe that this entire text is AI.
AI likelihood · overall
AIArticle text · 1,378 words · 1 segments analyzed
Benchmark the local LLMs you already have — speed, memory, and quality — as a live terminal leaderboard. homebench is a single-command TUI that discovers the models installed in your local runner (Ollama, LM Studio, llama.cpp, vLLM, or any OpenAI-compatible server), runs a curated quality suite, measures tokens/sec, time-to-first-token, and memory footprint on your actual machine, and renders a live comparison leaderboard. pip install homebench homebench That's it. No config, no API keys, no cloud. Why There are great tools for one half of this problem, but nothing local-first that does both: llama-bench (inside llama.cpp) measures speed only. lm-evaluation-harness measures quality but has no polished laptop UX and isn't built around the model runners most people actually use locally. homebench fills the gap: local-first, zero-config, UX-driven. Clone-and-run, point it at the models you already pulled, and get an at-a-glance answer to "which of my local models is actually good, and how fast is it on this laptop?" What it measures Metric How tok/s Output tokens ÷ generation time. Ollama reports server-side eval timing; OpenAI-compatible backends are timed client-side from the token stream. Excludes prompt processing and model load. TTFT Wall-clock time to the first streamed token (minus model-load time where the runner reports it). Memory Resident model size when the runner exposes it (Ollama /api/ps, LM Studio /api/v0), plus a best-effort peak-RSS sample of the backend's processes. Quality 31 deterministically-graded tasks across math, reasoning, factual recall, instruction-following/structured-output, extraction, and code understanding. Optional LLM-as-judge adds open-ended tasks (summaries, email, haiku, explanations). Install pip install homebench # then run: homebench Prefer an isolated install? Use pipx: pipx install homebench Or from source: git clone https://github.com/david-g-3654/homebench cd homebench pip install . Requires Python 3.9+. Usage homebench # fast default: 3 smallest models, quick suite (TUI) homebench --all # benchmark every discovered model homebench --full # run the full quality suite (not just the fast subset) homebench --no-tui # plain live renderer (great for piping / CI) homebench -m llama3.2,qwen3:8b # only these models homebench --limit 3 # cap the number of models homebench --provider lmstudio # use LM Studio instead of auto-detect homebench --provider llamacpp # llama.cpp server (llama-server) homebench --provider vllm # vLLM homebench --provider openai --host http://localhost:5000 # any OpenAI-compatible server homebench --refresh-cache # recompute instead of reusing cached responses homebench --no-quality # speed + memory only (fast) homebench --no-speed # quality only homebench --judge qwen3:8b # enable LLM-as-judge (adds open-ended tasks) homebench --tasks mypack.yaml # use a custom task pack instead of the built-in suite homebench --add-tasks mypack.yaml # add a pack on top of the built-in suite homebench --label "before tuning" # tag this run for later diffing homebench --md results.md # also export a Markdown report homebench --json results.json # also export raw JSON homebench list # just list discovered models homebench tasks # show the quality suite (add --tasks to preview a pack) homebench history # list past runs (saved automatically) homebench diff # diff the two most recent runs homebench diff 3 1 # diff run #3 (base) against run #1 (newer) homebench throughput # batch-throughput sweep (concurrency 1,2,4,8) homebench throughput --concurrency 1,8,16 --provider vllm homebench fit # which popular models fit YOUR hardware? Run homebench --help for the full flag list. Example output A real quick-suite run on an Apple M1 (16 GB), via Ollama: Final leaderboard ┏━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━━┳━━━━━━┳━━━━━━━┳━━━━━━━━┳━━━━━━━━┓ ┃ # ┃ Model ┃ Params ┃ Quality ┃ Pass ┃ tok/s ┃ TTFT ┃ Memory ┃ ┡━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━━╇━━━━━━╇━━━━━━━╇━━━━━━━━╇━━━━━━━━┩ │ 1 │ llama3.2:latest │ 3.2B │ 75% │ 6/8 │ 16.8 │ 545 ms │ 2.4 GB │ │ 2 │ alibayram/smollm3 │ 3.1B │ 38% │ 3/8 │ 16.9 │ 829 ms │ 2.1 GB │ └───┴──────────────────────┴────────┴─────────┴──────┴───────┴────────┴────────┘ (Numbers are for that laptop at that moment — see Limitations.) Providers At least one local model runner must be reachable: Provider --provider Default host Host env var Notes Ollama ollama http://localhost:11434 OLLAMA_HOST Native API; reports model memory via /api/ps. LM Studio lmstudio http://localhost:1234 LMSTUDIO_HOST Enriches metadata + memory via native /api/v0. llama.cpp llamacpp http://localhost:8080 LLAMACPP_HOST llama-server, OpenAI-compatible. vLLM vllm http://localhost:8000 VLLM_HOST Set VLLM_API_KEY if started with --api-key. OpenAI-compatible openai — OPENAI_BASE_URL Any /v1 server (Jan, LocalAI, TGI, …); pass --host. Auto-detection tries Ollama → LM Studio → llama.cpp → vLLM (the generic openai provider is explicit-only). Force one with --provider. Override host with --host or the env var above. How quality grading works The suite is small on purpose — enough tasks across categories to separate models, few enough that every model runs in a couple of minutes on a laptop. Each task is graded deterministically (exact numeric match, multiple-choice letter, substring, valid-JSON, regex). Temperature is 0 and a fixed seed is used for reproducibility. See homebench tasks for the list. The optional --judge MODEL flag turns on an LLM-as-judge (any local model) that scores open-ended tasks 1–5 against a reference answer. It's a signal, not an oracle. Fast by default Benchmarking every model on the full suite takes a while on a laptop, so the defaults are tuned for a quick first look: 3 smallest models by default (smallest first, so results appear fast) — --all for everything, -m to choose. A fast quality subset (~8 tasks across all categories) — --full for all 31. Response caching: quality runs use temperature 0 + a fixed seed, so responses are deterministic and cached under ~/.homebench. Re-running only regenerates new models/tasks (unchanged ones are re-graded from cache in milliseconds); --refresh-cache forces recompute, --no-cache disables it. In practice this turns a first run from ~15–25 min (all models, full suite) into ~1–2 min, and a re-run into seconds. For a thorough pass (CI, final numbers) use homebench --all --full. Custom task packs Bring your own evals with a JSON or YAML pack — no Python required. --tasks replaces the built-in suite; --add-tasks appends to it. YAML needs the optional extra (pip install "homebench[yaml]"); JSON works out of the box. # mypack.yaml — homebench --tasks mypack.yaml name: my-pack tasks: - id: capital_japan category: factual prompt: "What is the capital of Japan? Answer with just the city name." grader: {type: contains_any, values: ["Tokyo"]} reference: Tokyo - id: add category: math prompt: "What is 12 + 30? End with the answer on its own line." grader: {type: exact_number, value: 42} - id: explain # no grader -> open-ended, scored only with --judge category: open prompt: "Explain photosynthesis in one sentence." reference: "Plants convert sunlight, water, and CO2 into glucose and oxygen." Grader type values: exact_number (value, tol), multiple_choice (value), contains_any (values), regex (pattern, ignorecase), valid_json (keys), valid_json_array (length). Omit grader for a judge-only task. Runnable examples live in examples/; preview any pack with homebench tasks --tasks mypack.yaml. History & diffing Every run is saved automatically to $HOMEBENCH_HOME/runs (default ~/.homebench/runs); disable with --no-save, and tag runs with --label. homebench history # table of past runs (newest first) homebench diff # previous run -> latest homebench diff 3 # run #3 -> latest homebench diff 3 1 # run #3 (base) -> run #1 (newer) diff compares models by name and shows per-model deltas in quality and throughput, plus which models were added or removed between runs — handy for "did that quantization / setting actually help?" Batch throughput The main leaderboard measures single-stream tok/s. Servers that batch requests (vLLM, llama.cpp continuous batching, Ollama with OLLAMA_NUM_PARALLEL>1) can do far more total work under concurrency — homebench throughput measures that: homebench throughput -m my-model --concurrency 1,2,4,8 It fires N requests at each concurrency level (N defaults to 3×concurrency) and reports aggregate tok/s (total output ÷ wall-clock), the speedup vs. concurrency 1, mean per-request rate, and latency (mean / p95): Batch throughput — my-model (vllm) ┏━━━━━━┳━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━┓ ┃ Conc ┃ Reqs ┃ Agg tok/s ┃ Speedup ┃ Req tok/s ┃ Mean lat ┃ p95 lat ┃ Errors ┃ ┡━━━━━━╇━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━┩ │ 1 │ 4 │ 95.0 │ 1.00× │ 95.0 │ 1.35 s │ 1.4 s │ 0 │ │ 4 │ 12 │ 320.0 │ 3.37× │ 82.0 │ 1.56 s │ 1.9 s │ 0 │ │ 8 │ 24 │ 540.0 │ 5.68× │ 70.0 │ 1.83 s │ 2.6 s │ 0 │ └──────┴──────┴───────────┴─────────┴───────────┴──────────┴─────────┴────────┘ On a non-batching setup, aggregate throughput stays flat while latency climbs — which is itself a useful thing to see. Add --json FILE to export. What can my machine run?