20-task local benchmark

Can this local model build, debug, and recover a real system?

Compare local model variants, quantizations, and inference runtimes on 20 terminal tasks.

A self-contained Terminal-Bench 2.1 derivative; these are not official Terminal-Bench 2.1 scores.

results explorer

Compare runs

Loading committed benchmark results…

Methodology

How to read these results

Primary score: passed within attempts. A pass requires a final verifier reward of exactly 1. Pass@1 is the first-attempt view.

Sequential local runs. Harbor and Terminus use one OpenAI-compatible endpoint, with a three-hour agent timeout per attempt.

Full run identity. Each result retains the model, quant, platform, engine, backend, profile, tokens, timings, attempts, and transcripts.