Ox Alpha on DeepSWE

Score vs. efficiency on the same deterministic 10-task DeepSWE subset (seed 0). Leaderboard models are the mean of 4 attempts per task from DeepSWE's public trial data; ox-alpha is a single local run with the official harness.

ox-alpha (1 run) leaderboard models (mean of 4 runs, same tasks)

Score vs. agent steps

Avg agent steps per task — fewer is more efficient ↗

Score vs. output tokens

Avg output tokens per task — fewer is more efficient ↗

Underlying values (10-task subset)
ModelEffortScoreAvg stepsAvg output tokensAttempts/task

Ox-alpha's score carries single-run noise (±1 task = ±10 points); leaderboard values average 4 attempts per task. Steps and tokens for leaderboard models come from DeepSWE's per-trial public data (claude-fable-5: 38 of 40 trials had parseable metrics); ox-alpha's from local mini-swe-agent trajectories with exact API-reported token usage. All runs use the same agent (mini-swe-agent) and task subset.