Terminal-Bench 4 subset methodology
We run 19 of Terminal-Bench 4.0's 66 tasks, five attempts each. A full run is 330 trials and costs thousands of dollars; the subset is 95 trials at about 15% of the cost. This page explains which tasks we picked, why, and how far a subset score can stand in for a full one.
Which tasks, and why
The subset has to run where we run: HF Jobs, one container per trial, no GPU. Three rules take the 66 tasks down to an eligible pool of 48:
| Rule | Tasks excluded |
|---|---|
| Single container: no docker-compose sidecar | ctr-optimization, cumulative-layout-shift, freight-dispatch-shift, heat-pump-warranty, intrastat-meldung, kv-live-surgery, legacy-utility-triage, live-database-cutover, medical-claims-processing, nextjs-performance, payments-pipeline-fix |
| No GPU | fp8-rmsnorm-gemm, jax-speedrun-gpu, math-eval-grader |
| No safety refusals: no leaderboard trial ended in a model refusal | batched-eval-parity, interleaved-vigenere, kv-live-surgery, shadow-relay, uefi-bootkit |
From that pool we chose 18 tasks that, together, track the full leaderboard closely at low cost, then added hof-topology-interpenetration. Four other candidates (distributed-dedup, ks-solver-cpp, vf2-speedup-networkx, embedding-drift-monitor) were rejected because their verifiers have open defect reports.
The 19 tasks: atrx-vep-crispr, cad-model, cargo-flight-dispatch, fin-saccr-rwa, gsea-proteomics, hof-topology-interpenetration, html-js-filter, mvcc-lsm-compaction, photonic-waveguide-routing, production-planning, react-lead-form, satb-audio-transcription, session-window-debug, sound-change-cascade, telecom-entity-resolution, vba-userform-port, vllm-deepseek-streaming, vpp-loss-divergence and wal-recovery-ordering.
How representative is it?
We scored every Terminal-Bench 4.0 leaderboard row on the 19 tasks alone and compared it with the row's published full score. 26 of the 27 rows qualify; Opus 5 · max is left out because its trial records (173 passes) don't reproduce its published score (171).
Loading chart…
| 19-task subset | Random 19 from the eligible 48 (20,000 draws) | |
|---|---|---|
| Average gap to the full score | 3.25 points | median 5.16; only 6% of draws are closer |
| Within ±5 / ±10 points | 22 / 26 · 25 / 26 rows | |
| Rank correlation with the full score (Spearman) | 0.93 | median 0.94; 65% of draws rank better |
| Against the other 47 tasks only (held out) | 4.56 points, Spearman 0.90 | |
| Cost of a subset run, as a share of a full run | 15% (wall time 16%) | median 24%; under 1% of draws are as cheap |
The subset is closer to the full score than almost any random pick of the same size, and much cheaper, but random picks rank runs slightly better: it was chosen for score and cost, not ranking. The held-out row is the stricter test, because the 19 tasks are part of the full score they're compared with.
Estimating full-run cost
On the 15 rows with complete cost records, a full run costs a median 6.4× the subset (5.3× to 8.4×). Using that multiplier to estimate a row's full-run cost is off by 9.7% on average (worst 23.9%). Scaling by task count (66 ÷ 19) underestimates every row, by about half: the subset's tasks are cheaper than average.
Caveats
- Chosen with this leaderboard in view. The agreement above flatters the subset relative to a model it has never seen.
- Rows share models. The 26 rows cover 15 models, so they aren't independent samples.
- Wide error bars. With 19 tasks, a single run's standard error is about ±10 points. Treat gaps under ten points as unresolved.
- The range is slightly compressed. On the weakest third of rows the subset reads 2.3 points high on average; on the strongest third, 0.6 points low.
- Not every task separates models. cargo-flight-dispatch hasn't been passed by any leaderboard row.
- Task revisions. Six rows (GPT-6 Astra × 5 and Gemini 3.8 Flash) ran an earlier revision of some subset tasks than the other 21.
Costs on the page
Leaderboard rows show the sum of their recorded trial costs on the 19 tasks; where a task has a trial without a recorded cost, the figure is marked ~. Our subset runs show Harbor's recorded token cost. Each comparator's run page also shows its full 66-task run and published cost.