Score breakdown
Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.
Why it matters
Useful for local-AI builders who want a single command that tells them which of the models they already pulled is the best fit for their laptop, with a reproducible quality suite and real speed/memory numbers rather than a vibes-based pick. The 31 deterministically-graded tasks, the auto-detection of Ollama / LM Studio / llama.cpp / vLLM / MLX, the response cache that turns a re-run into seconds
Where this stands now
homebench ranks #728 of 3231 tracked Radar items by composite score (7.8 against a section median of 4.9). The section currently carries 2115 Bronze, 645 Gold, 471 Silver. RepoRadar has retained observations for this record since 2026-08-03 (68 days in the current window). Signal extremes versus the section: momentum at the 81th percentile; novelty at the 77th percentile.
Who should use it
Who should skip it
Pass on homebench if its scope or audience does not match what your team is building right now.
About this signal
homebench is tracked by RepoRadar as a code repository in the Radar section. First seen 2026-08-03; the source record was last checked on 2026-08-03. The current verdict is 'try now' with a Gold tier and easy setup difficulty. Across RepoRadar's eight signals, homebench is strongest on workflow potential (8.9) and open-source/build quality (8.4) and weakest on evidence quality (7.2) — a profile worth weighing against your own priorities. This page summarizes the evidence RepoRadar captured from https://github.com/david-g-3654/homebench.
How this item is evaluated
The homebench record combines a 7.8/10 composite score with separate popularity (51.0), risk (none), and setup (easy) signals. See the scoring methodology for the current weights and evidence definitions.
Questions worth asking before you adopt this
Putting this into practice? Read How to evaluate an AI tool before you adopt it for the checklist behind this score.
Risk explanation
the quality suite is intentionally small (31 tasks) and the README is explicit that the homebench value score is a laptop-fit signal within a single run, not a global benchmark; do not compare homebench value scores across different laptops; response caching uses temperature 0 + a fixed seed and reuses cached responses under ~/.homebench; pass --refresh-cache after changing a model file or after upgrading a runner version to avoid stale results; the optional --judge MODEL flag adds LLM-as-judge for open-ended tasks; the README explicitly says the judge score 'is a signal, not an oracle'; the MLX provider shares llama.cpp's default port (8080); homebench detects MLX only when --provider mlx is passed explicitly.