Benchmarks/Agent comparison

We gave every agent the same bugs

freecode against other coding agents on SWE-bench Lite django/django. Same tasks, same model, same key, and every agent at full autonomy.

runs
1, latest 2026-09-06T07-18-04-242Z
instances
10 shared
graded
SWE-bench harness
isolation
container

Run history

Every row on this page comes from this single run — nothing was merged in.

  1. #12026-09-06 10:12 UTCfreecode ·claude-codelatest2026-09-06T07-18-04-242Z
77% vs 80%
resolved
freecode vs claude-code, over 10 shared instances
60% fewer
tokens per trial
freecode 287k vs claude-code 710k
1.2× slower
median wall time
freecode 76s vs claude-code 64s
2.8× cheaper
cost per trial
freecode $0.0262 vs claude-code $0.0728

Setup

The autonomy flag is part of the experiment, not a footnote: running one agent at full autonomy against another at its default measures permission defaults rather than agents.

agentversionmodelautonomy
freecode0.30.1minimax/MiniMax-M3--agent danger: no permission prompts
claude-code2.1.251 (Claude Code)MiniMax-M3--dangerously-skip-permissions: all permission checks bypassed

Resolved

Marked resolved by the official SWE-bench grader. Computed over the 10 instances every agent in this matchup ran.

freecode77%
claude-code80%
freecode
23/30
claude-code
24/30

Median wall time

Spawn to agent exit. Shorter is better — the longest bar is the slowest agent, not the winner. An agent that runs the tests pays for it here, which is not obviously a vice.

freecode76s
claude-code64s

Where the time goes

freecode~20 turns per trial·~3.8s per turn
claude-code~20 turns per trial·~3.2s per turn

Wall time is turns × time per turn: an agent can lose this bar by making slow turns, or by making more of them. The split above says which.

Slowest median in this matchup: 76s.

Median patch size

Bytes of diff produced. A fact about each harness's style, not a score — a bigger patch is not a better fix, and a smaller one is not automatically surgical.

freecode659B
claude-code642B

New files created

None — across all 60 trials, every patch edited existing code only. That is what a bug-fix task set should look like: a harness that scaffolds new files here would be a smell, not a feature.

Tokens per trial

Input + output counted off one recording proxy for every agent — the same meter, the same rules, never each vendor's own accounting. Fewer is cheaper, but an agent that gives up early is also cheap; read this next to the outcome bar, not instead of it.

freecode287k
claude-code710k

Prompt cache

How much of each trial's input the provider served from its prompt cache, off the same meter as the token chart. The pale segment was read from cache at a discounted rate; the solid segment is fresh context, billed in full. Bars share one scale, so their lengths also compare input volume.

freecode94.4% cached
claude-code91.6% cached
read from cache (discounted)fresh input (billed in full)

A high hit rate means the harness keeps its prompt stable enough for the provider to reuse it. What each trial actually pays full price for is the solid segment: freecode ~16k fresh input tokens per trial vs claude-code ~59k — the hit rate is where a token gap becomes a cost gap.

Cost per trial

USD from a committed rate card applied to the proxy's counts — identical pricing for every agent. Cache reads are discounted, not added. The worst-single-trial figure is there because a good mean can hide one runaway.

freecode$0.0262
claude-code$0.0728

On the 8 bugs every agent solved: freecode $0.0138 (147k tok) · claude-code $0.0265 (257k tok). Averaging over failures would make the quitter cheapest — this mean covers solved bugs only.

Cost distribution

Every metered trial, one dot, by bug. A mean hides the outliers — this does not. An open ring is a trial that ran but did not resolve, so a cheap dot low on the chart is not always a win.

$0.000$0.122$0.244$0.366$0.48810914109241100111019110391104911099111331117911283freecode · 10914 · $0.0132 · resolvedclaude-code · 10914 · $0.0181 · resolvedfreecode · 10914 · $0.0099 · resolvedclaude-code · 10914 · $0.0551 · resolvedfreecode · 10914 · $0.0356 · resolvedclaude-code · 10914 · $0.0278 · resolvedfreecode · 10924 · $0.0257 · resolvedclaude-code · 10924 · $0.0398 · resolvedfreecode · 10924 · $0.0327 · resolvedclaude-code · 10924 · $0.0271 · resolvedfreecode · 10924 · $0.0319 · not resolvedclaude-code · 10924 · $0.2999 · not resolvedfreecode · 11001 · $0.0232 · resolvedclaude-code · 11001 · $0.0283 · resolvedfreecode · 11001 · $0.0178 · resolvedclaude-code · 11001 · $0.0847 · resolvedfreecode · 11001 · $0.0150 · resolvedclaude-code · 11001 · $0.0335 · resolvedfreecode · 11019 · $0.1227 · not resolvedclaude-code · 11019 · $0.4438 · not resolvedfreecode · 11019 · $0.0665 · not resolvedclaude-code · 11019 · $0.2590 · not resolvedfreecode · 11019 · $0.1252 · not resolvedclaude-code · 11019 · $0.2644 · not resolvedfreecode · 11039 · $0.0068 · resolvedclaude-code · 11039 · $0.0185 · resolvedfreecode · 11039 · $0.0110 · resolvedclaude-code · 11039 · $0.0264 · resolvedfreecode · 11039 · $0.0116 · resolvedclaude-code · 11039 · $0.0167 · resolvedfreecode · 11049 · $0.0143 · resolvedclaude-code · 11049 · $0.0241 · resolvedfreecode · 11049 · $0.0169 · resolvedclaude-code · 11049 · $0.0074 · resolvedfreecode · 11049 · $0.0131 · resolvedclaude-code · 11049 · $0.0301 · resolvedfreecode · 11099 · $0.0036 · resolvedclaude-code · 11099 · $0.0106 · resolvedfreecode · 11099 · $0.0016 · resolvedclaude-code · 11099 · $0.0104 · resolvedfreecode · 11099 · $0.0016 · resolvedclaude-code · 11099 · $0.0150 · resolvedfreecode · 11133 · $0.0182 · resolvedclaude-code · 11133 · $0.0213 · resolvedfreecode · 11133 · $0.0137 · resolvedclaude-code · 11133 · $0.0361 · resolvedfreecode · 11133 · $0.0172 · resolvedclaude-code · 11133 · $0.0311 · resolvedfreecode · 11179 · $0.0050 · resolvedclaude-code · 11179 · $0.0118 · resolvedfreecode · 11179 · $0.0028 · resolvedclaude-code · 11179 · $0.0122 · resolvedfreecode · 11179 · $0.0061 · resolvedclaude-code · 11179 · $0.0243 · resolvedfreecode · 11283 · $0.0329 · not resolvedclaude-code · 11283 · $0.0576 · not resolvedfreecode · 11283 · $0.0417 · not resolvedclaude-code · 11283 · $0.1914 · not resolvedfreecode · 11283 · $0.0472 · not resolvedclaude-code · 11283 · $0.0586 · resolved
freecodeclaude-codeopen ring = not resolved

What each fix cost

Only the 8bugs every agent resolved (spec §7.2) — same bug, same model, so the gap is how much context each harness moved to get there. Cost is the mean over that agent's resolved trials of the bug; the cheaper agent on each row is in bold.

Set default FILE_UPLOAD_PERMISSION to 0o644.django-10914
1.7× cheaper1.9× fewer tokens
freecode
$0.0196 (205k)
claude-code
$0.0336 (383k)
Allow FilePathField path to accept a callable.django-10924
1.1× cheaper1.1× more tokens
freecode
$0.0292 (329k)
claude-code
$0.0334 (287k)
Incorrect removal of order_by clause created as multiline RawSQLdjango-11001
2.6× cheaper2.0× fewer tokens
freecode
$0.0187 (213k)
claude-code
$0.0488 (429k)
sqlmigrate wraps it's outpout in BEGIN/COMMIT even if the database doesn't support transactional DDLdjango-11039
2.1× cheaper1.9× fewer tokens
freecode
$0.0098 (101k)
claude-code
$0.0206 (193k)
Correct expected format in invalid DurationField error messagedjango-11049
1.4× cheaper1.3× fewer tokens
freecode
$0.0147 (158k)
claude-code
$0.0206 (213k)
UsernameValidator allows trailing newline in usernamesdjango-11099
5.2× cheaper4.8× fewer tokens
freecode
$0.0023 (21k)
claude-code
$0.0120 (102k)
HttpResponse doesn't handle memoryview objectsdjango-11133
1.8× cheaper2.0× fewer tokens
freecode
$0.0164 (174k)
claude-code
$0.0295 (354k)
delete() on instances of models without any dependencies doesn't clear PKs.django-11179
3.5× cheaper2.8× fewer tokens
freecode
$0.0046 (38k)
claude-code
$0.0161 (108k)
mean over 8 bugs
1.9× cheaper1.7× fewer tokens
freecode
$0.0138 (147k)
claude-code
$0.0265 (257k)

Per instance

One row per bug, one cell per trial. Disagreement between trials of the same agent is the interesting signal — it is what a single-run benchmark cannot show you.

resolvedpatched, not resolvedno patch
Set default FILE_UPLOAD_PERMISSION to 0o644.django__django-10914
freecodet1 · 52sfreecodet2 · 36sfreecodet3 · 338sclaude-codet1 · 20sclaude-codet2 · 164sclaude-codet3 · 150s
Allow FilePathField path to accept a callable.django__django-10924
freecodet1 · 189sfreecodet2 · 116sfreecodet3 · 134sclaude-codet1 · 412sclaude-codet2 · 68sclaude-codet3 · 516s
Incorrect removal of order_by clause created as multiline RawSQLdjango__django-11001
freecodet1 · 86sfreecodet2 · 76sfreecodet3 · 77sclaude-codet1 · 71sclaude-codet2 · 214sclaude-codet3 · 92s
Merging 3 or more media objects can throw unnecessary MediaOrderConflictWarningsdjango__django-11019
freecodet1 · 410sfreecodet2 · 211sfreecodet3 · 528sclaude-codet1 · 841sclaude-codet2 · 900sclaude-codet3 · 900s
sqlmigrate wraps it's outpout in BEGIN/COMMIT even if the database doesn't support transactional DDLdjango__django-11039
freecodet1 · 35sfreecodet2 · 65sfreecodet3 · 53sclaude-codet1 · 17sclaude-codet2 · 27sclaude-codet3 · 47s
Correct expected format in invalid DurationField error messagedjango__django-11049
freecodet1 · 74sfreecodet2 · 91sfreecodet3 · 81sclaude-codet1 · 30sclaude-codet2 · 13sclaude-codet3 · 43s
UsernameValidator allows trailing newline in usernamesdjango__django-11099
freecodet1 · 12sfreecodet2 · 10sfreecodet3 · 11sclaude-codet1 · 10sclaude-codet2 · 11sclaude-codet3 · 18s
HttpResponse doesn't handle memoryview objectsdjango__django-11133
freecodet1 · 83sfreecodet2 · 38sfreecodet3 · 64sclaude-codet1 · 45sclaude-codet2 · 60sclaude-codet3 · 79s
delete() on instances of models without any dependencies doesn't clear PKs.django__django-11179
freecodet1 · 20sfreecodet2 · 16sfreecodet3 · 16sclaude-codet1 · 12sclaude-codet2 · 12sclaude-codet3 · 51s
Migration auth.0011_update_proxy_permissions fails for models recreated as a proxy.django__django-11283
freecodet1 · 144sfreecodet2 · 156sfreecodet3 · 146sclaude-codet1 · 136sclaude-codet2 · 304sclaude-codet3 · 153s

What this does not tell you

Published rather than managed. A benchmark whose limitations live in a footnote is an advertisement.

10 instances is a demo, not a leaderboard

SWE-bench Lite is 300 instances across 11 repositories. One repository's idioms are not the field, and a handful of instances from it is an anecdote with decimal places.

Contamination is unfixable on this task set

Every SWE-bench Lite fix is public and predates the training cutoff of the models involved. Closing the network stops lookup, not recall. That is survivable for a relative comparison — every agent gets the same unfair advantage — but it invalidates any absolute reading.

This compares harnesses, not models

Every agent is pinned to the same model (MiniMax-M3) on the same key, and each keeps its own system prompt. Change the model and it becomes a different experiment with the same table.

We built the harness and we are in the table

Every incentive here points one way. The countermeasures are structural: the task set is external, the grader is external, the artifacts are published, and a loss goes in the headline.

How it ran

The whole pipeline, in order. Every step is the same for every agent — the moment one step differs per agent, the benchmark stops comparing harnesses and starts comparing our treatment of them.

  1. 1

    Pick the bugs

    10 real django/django issues from SWE-bench Lite, fetched from HuggingFace. The gold patch, the test patch, and the maintainer hints are stripped before anything touches disk — the answer key never enters this repo.

  2. 2

    Build the workspace

    Each trial gets its own fresh checkout at the commit just before the real fix landed, cloned from a cached local mirror on a detached HEAD. No network clone in the timed path, nothing shared between trials.

  3. 3

    Give every agent the same prompt

    One string, identical for all agents: the upstream issue text, plus instructions not to touch tests and not to commit. The prompt lives in its own file so changing it reviews as what it is — a change to the experiment.

  4. 4

    Pin the model, max the autonomy

    Every agent runs MiniMax-M3 on the same API key, each with its own full-autonomy flag (see Setup above). Running one agent at full autonomy against another at its default would measure permission defaults, not agents.

  5. 5

    Isolate each trial in Docker

    Each trial runs in a container on an internal Docker network whose only exit is the metering proxy — the agent cannot look the fix up, and a fresh $HOME means nothing an agent learns carries into the next trial.

  6. 6

    Meter everything through one proxy

    All model traffic passes through a single recording pass-through proxy. Tokens are counted by that one meter and priced from a committed rate card — the same accounting for every agent, never each vendor's own dashboard.

  7. 7

    Run the matrix, keep the diff

    agents × bugs × trials, 3 trials per bug here. When the agent exits, its working-tree diff is extracted as the patch — that diff is the entire submission.

  8. 8

    Grade with the official harness

    The official SWE-bench harness applies each patch in its own Docker environment and runs the project's real test suite. Only its verdict marks a trial resolved — 'produced a patch' is never promoted to 'fixed the bug'.

  9. 9

    Publish the evidence

    Every trial writes its prompt, argv, patch, and stdout/stderr to the results directory, and the numbers land on this page by finishing. There is no editorial step between the run and what you are reading.

Check it yourself

The harness, the agent adapters, and the task list are all in the repo. A run lands on this page by finishing — there is no editorial step between the numbers and you.

pnpm bench:agents --agents freecode,claude-code --trials 3

Every trial writes its full evidence — prompt, argv, patch, stdout/stderr — to bench/agent-bench/results/<run>/. The task set is SWE-bench Lite, fetched from HuggingFace with the answer key stripped before anything touches disk.