20-task local benchmark
Can this local model build, debug, and recover a real system?
Compare local model variants, quantizations, and inference runtimes on 20 terminal tasks.
A self-contained Terminal-Bench 2.1 derivative; these are not official Terminal-Bench 2.1 scores.
results explorer
Compare runs
Model runs
Select up to four columns for direct comparison.
Run comparison
Baseline and recovery metrics for the selected deployments.
Task matrix
Open any result to see attempts, failure classification, verifier evidence, and transcripts.
Methodology
How to read these results
Primary score: passed within attempts. A pass requires a final verifier reward of exactly 1. Pass@1 is the first-attempt view.
Sequential local runs. Harbor and Terminus use one OpenAI-compatible endpoint, with a three-hour agent timeout per attempt.
Full run identity. Each result retains the model, quant, platform, engine, backend, profile, tokens, timings, attempts, and transcripts.