Item detail
github.com

homebench

homebench is a code repository in RepoRadar's Radar section, holding Gold tier and a 'try now' verdict. Its strongest signal is workflow potential, scored 8.9 out of 10.

Score7.8
Popularity51.0
Risknone
TierGold
Score breakdown
Usefulness8.2
Novelty7.4
Momentum7.8
Maturity7.2
Open-source/build8.4
Evidence7.2
Workflow potential8.9
Setup ease8.2

Popularity is tracked separately. Support, ads, sponsorships, and tips never affect these signals.

Why it matters

Useful for local-AI builders who want a single command that tells them which of the models they already pulled is the best fit for their laptop, with a reproducible quality suite and real speed/memory numbers rather than a vibes-based pick. The 31 deterministically-graded tasks, the auto-detection of Ollama / LM Studio / llama.cpp / vLLM / MLX, the response cache that turns a re-run into seconds

Where this stands now

homebench ranks #728 of 3231 tracked Radar items by composite score (7.8 against a section median of 4.9). The section currently carries 2115 Bronze, 645 Gold, 471 Silver. RepoRadar has retained observations for this record since 2026-08-03 (68 days in the current window). Signal extremes versus the section: momentum at the 81th percentile; novelty at the 77th percentile.

Who should use it

local-AI builders who want a one-command answer to 'which of the models I already pulled is the best fit for my laptop' developers iterating on a quantization setting or prompt change who need a reproducible speed + quality + memory baseline to compare against the next run CI pipelines that want to gate a prompt or quantization change on a --fail-on-regression check that exits non-zero when quality or speed drops past a threshold

Who should skip it

Pass on homebench if its scope or audience does not match what your team is building right now.

About this signal

homebench is tracked by RepoRadar as a code repository in the Radar section. First seen 2026-08-03; the source record was last checked on 2026-08-03. The current verdict is 'try now' with a Gold tier and easy setup difficulty. Across RepoRadar's eight signals, homebench is strongest on workflow potential (8.9) and open-source/build quality (8.4) and weakest on evidence quality (7.2) — a profile worth weighing against your own priorities. This page summarizes the evidence RepoRadar captured from https://github.com/david-g-3654/homebench.

How this item is evaluated

The homebench record combines a 7.8/10 composite score with separate popularity (51.0), risk (none), and setup (easy) signals. See the scoring methodology for the current weights and evidence definitions.

Questions worth asking before you adopt this

Putting this into practice? Read How to evaluate an AI tool before you adopt it for the checklist behind this score.

Risk explanation

the quality suite is intentionally small (31 tasks) and the README is explicit that the homebench value score is a laptop-fit signal within a single run, not a global benchmark; do not compare homebench value scores across different laptops; response caching uses temperature 0 + a fixed seed and reuses cached responses under ~/.homebench; pass --refresh-cache after changing a model file or after upgrading a runner version to avoid stale results; the optional --judge MODEL flag adds LLM-as-judge for open-ended tasks; the README explicitly says the judge score 'is a signal, not an oracle'; the MLX provider shares llama.cpp's default port (8080); homebench detects MLX only when --provider mlx is passed explicitly.

Evidence links
Closest alternatives / related signals
local-llm benchmark tui ollama lm-studio llama.cpp vllm mlx
Verification record

What RepoRadar actually verified

Discovered

Automated discovery and source capture. Last checked 2026-10-10T17:09:29.745535Z.

No editorial or hands-on review is claimed. This record remains at Discovered.

Verification sources

Longitudinal intelligence

How this decision record is moving

Raw history JSON →

50 dated snapshots retained from 2026-08-09 through 2026-10-10; see the snapshot index for explicit coverage gaps. Stars, version, release, pricing, integration, risk, maintenance, verdict, score, and momentum fields remain explicit even when a source has not reported them. Repository momentum is a normalized 0–10 RepoRadar signal; GitHub stars appear only where the popularity monitor retained exact timestamped observations.

RepoRadar score7.8 current · +0.0 net
Repository momentum5.5 current · -3.5 net
GitHub stars (observed)58 current · +7 net
GitHub stars58 exact observation
Versionv0.11.0
Last release2026-08-10T03:05:17Z
Maintenancemaintained
Current riskNone
Current verdictTry now
Pricing baselineNo structured commercial pricing baseline
Pricing checkedNot applicable or not recorded
Pricing freshnessNo dated commercial pricing review
Integrations baselineOllama, vLLM

Recent dated points

DateScoreMomentumStarsRiskVerdictMaintenance
2026-10-107.85.558NoneTry nowmaintained
2026-10-097.85.558NoneTry nowmaintained
2026-10-087.85.858NoneTry nowmaintained
2026-10-077.85.557NoneTry nowmaintained
2026-10-067.85.557NoneTry nowmaintained
2026-10-057.85.557NoneTry nowmaintained
2026-10-037.85.557NoneTry nowmaintained
2026-10-027.85.557NoneTry nowmaintained
2026-10-017.85.557NoneTry nowmaintained
2026-09-307.85.557NoneTry nowmaintained
2026-09-297.85.557NoneTry nowmaintained
2026-09-287.85.557NoneTry nowmaintained

Why the record changed

Stars change

Stars changed: 57 → 58.

Stars change

Stars changed: 56 → 57.

Maintenance change

Maintenance changed: Active → maintained.

Stars change

Stars changed: 55 → 56.

Stars change

Stars changed: 54 → 55.

Stars change

Stars changed: 53 → 54.

Stars change

Stars changed: 51 → 53.

Version change

Version changed: v0.10.0 → v0.11.0.

Version change

Version changed: not recorded → v0.10.0.

Maintenance change

Maintenance changed: Not yet measured → Active.