Score vs. efficiency pooled over two disjoint deterministic DeepSWE subsets — 20 of the 113 tasks. Leaderboard models are the mean of 4 attempts per task from DeepSWE's public trial data; ox-alpha is a single local run with the official harness.
Avg agent steps per task — fewer is more efficient ↗
Avg output tokens per task — fewer is more efficient ↗
| Model | Effort | Score | Avg steps | Avg output tokens | Attempts/task |
|---|
Ox-alpha's score carries single-run noise (±1 task = ±5 points on 20 tasks); batch 1 scored 8/10, batch 2 scored 5/10; leaderboard values average 4 attempts per task. Steps and tokens for leaderboard models come from DeepSWE's per-trial public data (claude-fable-5: 38 of 40 trials had parseable metrics); ox-alpha's from local mini-swe-agent trajectories with exact API-reported token usage. All runs use the same agent (mini-swe-agent) and task subset.