Ox Alpha on DeepSWE

Score vs. efficiency pooled over two disjoint deterministic DeepSWE subsets — 20 of the 113 tasks. Leaderboard models are the mean of 4 attempts per task from DeepSWE's public trial data; ox-alpha is a single local run with the official harness.

ox-alpha (1 attempt/task, 20 tasks) leaderboard models (mean of 4 attempts/task, same 20 tasks)

Score vs. agent steps

Avg agent steps per task — fewer is more efficient ↗

Score vs. output tokens

Avg output tokens per task — fewer is more efficient ↗

Underlying values (pooled 20-task subsets)
ModelEffortScoreAvg stepsAvg output tokensAttempts/task

Ox-alpha's score carries single-run noise (±1 task = ±5 points on 20 tasks); batch 1 scored 8/10, batch 2 scored 5/10; leaderboard values average 4 attempts per task. Steps and tokens for leaderboard models come from DeepSWE's per-trial public data (claude-fable-5: 38 of 40 trials had parseable metrics); ox-alpha's from local mini-swe-agent trajectories with exact API-reported token usage. All runs use the same agent (mini-swe-agent) and task subset.