LEB-100-A v1.1 · Support-ticket panel of an internet provider
The support-ticket panel of an internet provider, written in 2013-style PHP: data-access functions and an index.php that routes, authorizes and builds the HTML. About 300 lines on PHP 8, mysqli and MySQL 8, with 13 planted flaws and 2 decoys.
-
1
Claude Sonnet 5.5 Anthropic · effort xhigh 3 of 3 runs (825 · 809 · 724)809 of 1000 LEB Gold
- Security 250/250
- Architecture 25/200
- Bugs 139/150
- Performance 150/150
- Clean code 100/100
- Compatibility 100/100
- Explanation 45/50
Comment and details
The most complete single-agent result: security at 250 of 250 in its first two runs, performance at full marks, compatibility untouched, and the CSV formula injection fixed the way the answer key expects. Its third run left three security flaws found but unfixed and fell to 724; like every agent, it left the dispatcher whole.
- Full marks
- Security, Performance, Clean code, Compatibility
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority
Run by run
- Run 1 · 825 11 of 13 flaws fixed no false positive contract kept 22/22 checks 19min US$ 3.60 Scorecard
- Run 2 · 809 11 of 13 flaws fixed no false positive contract kept 22/22 checks 23min US$ 2.96 Scorecard
- Run 3 · 724 8 of 13 flaws fixed no false positive contract kept 22/22 checks 23min US$ 2.91 Scorecard
-
2
Claude Sonnet 5.5 Anthropic · effort max 3 of 3 runs (807 · 773 · 820)807 of 1000 LEB Gold
- Security 250/250
- Architecture 12/200
- Bugs 150/150
- Performance 150/150
- Clean code 100/100
- Compatibility 100/100
- Explanation 45/50
Comment and details
The most consistent of the top three: 773 to 820 across three runs, bugs and performance at full marks in all of them, no false positive and no contract broken. More effort than xhigh bought nothing measurable, at about one and a half times the cost.
- Full marks
- Security, Bugs, Performance, Clean code, Compatibility
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority
Run by run
- Run 1 · 807 11 of 13 flaws fixed no false positive contract kept 22/22 checks 27min US$ 3.86 Scorecard
- Run 2 · 773 10 of 13 flaws fixed no false positive contract kept 22/22 checks 29min US$ 4.75 Scorecard
- Run 3 · 820 11 of 13 flaws fixed no false positive contract kept 22/22 checks 41min US$ 6.65 Scorecard
-
3
Claude Sonnet 5.5 Anthropic · effort max (ultracode) 3 of 3 runs (774 · 820 · 759)774 of 1000 LEB Gold
- Security 250/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 100/100
- Compatibility 100/100
- Explanation 45/50
Comment and details
Claude Code's multi-agent mode: 68 to 145 subagents and 2h to 7h 55min a run, at US$ 119 to 233 each, for 774, 820 and 759. At best it tied the same model working alone in half an hour, with the best-calibrated report of the Claude agents and the same untouched architecture.
- Full marks
- Security, Performance, Clean code, Compatibility
- Zero
- Architecture
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority
Run by run
- Run 1 · 774 11 of 13 flaws fixed no false positive contract kept 22/22 checks 7h 55min US$ 233.08 Scorecard
- Run 2 · 820 11 of 13 flaws fixed no false positive contract kept 22/22 checks 2h US$ 171.18 Scorecard
- Run 3 · 759 10 of 13 flaws fixed no false positive contract kept 22/22 checks 3h 34min US$ 119.07 Scorecard
-
4
Claude Fable 5.1 Anthropic · effort xhigh 2 of 3 runs (781 · 764) · not official764 of 1000 LEB Gold
- Security 233/250
- Architecture 0/200
- Bugs 139/150
- Performance 150/150
- Clean code 100/100
- Compatibility 100/100
- Explanation 42/50
Comment and details
Finds as much as Sonnet (11 to 12 of the 13 flaws) and fixes less of it: it explained the CSV formula injection and left it in place on purpose, to protect the file's consumers. Two runs so far; its client switched part of a third attempt to another model, which voided it.
- Full marks
- Performance, Clean code, Compatibility
- Zero
- Architecture
- Never fixed, in any run
- Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority
Run by run
-
5
Claude Opus 5.5 Anthropic · effort xhigh 3 of 3 runs (711 · 717 · 805)717 of 1000 LEB Silver
- Security 233/250
- Architecture 50/200
- Bugs 139/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 45/50
Comment and details
The highest architecture score in an official run (50 of 200), and the widest spread among the Claude models: 711 and 717, then 805. It fixed 9 flaws in every run and never flattened the nested ifs in the run that counts, which is where its gap to Sonnet comes from.
- Full marks
- Performance, Compatibility
- Zero
- Clean code
- Never fixed, in any run
- Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority
Run by run
- Run 1 · 711 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 16min Scorecard
- Run 2 · 717 9 of 13 flaws fixed no false positive contract kept 22/22 checks 18min US$ 3.01 Scorecard
- Run 3 · 805 9 of 13 flaws fixed no false positive contract kept 22/22 checks 15min US$ 2.71 Scorecard
-
6
GPT-6-astra OpenAI · effort ultra 3 of 3 runs (668 · 666 · 651)666 of 1000 LEB Silver
- Security 224/250
- Architecture 25/200
- Bugs 129/150
- Performance 150/150
- Clean code 25/100
- Compatibility 70/100
- Explanation 43/50
Comment and details
The strongest GPT agent, with two or three subagents in 10 to 15 minutes a run and three runs within 17 points of each other. It changed one business value in every run (−30 each), and its reports are among the best calibrated here (Brier 0.000).
- Full marks
- Performance
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 668 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 11min Scorecard
- Run 2 · 666 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 15min Scorecard
- Run 3 · 651 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 10min Scorecard
-
7
GPT-6.1-sol OpenAI · effort xhigh 3 of 3 runs (666 · 653 · 661)661 of 1000 LEB Silver
- Security 246/250
- Architecture 25/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 41/50
Comment and details
Security at 246 of 250, the best outside Claude, with the CSV formula injection and MD5 both fixed in the official run. What holds it below the Claude models is one business value changed in each run (−30) and the nested ifs left as they were.
- Full marks
- Performance
- Zero
- Clean code
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 666 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 27min Scorecard
- Run 2 · 653 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 26min Scorecard
- Run 3 · 661 10 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 24min Scorecard
-
8
GPT-6.1-sol pro OpenAI · effort xhigh 1 of 3 runs · not official no published cutoff or earlier release654 of 1000 LEB Silver
- Security 228/250
- Architecture 0/200
- Bugs 139/150
- Performance 150/150
- Clean code 25/100
- Compatibility 70/100
- Explanation 42/50
Comment and details
One run so far, through opencode, at US$ 1.01: 654, close to GPT-6.1-sol in Codex CLI. It left the CSV formula injection in place and changed one business value (−30).
- Full marks
- Performance
- Zero
- Architecture
- Never fixed, in any run
- Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 654 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 16min US$ 1.01 Scorecard
-
9
Grok 4.7 xAI · effort high 3 of 3 runs (638 · 663 · 607)638 of 1000 LEB Silver
- Security 207/250
- Architecture 0/200
- Bugs 139/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 42/50
Comment and details
The strongest model from outside Anthropic and OpenAI, with compatibility at 100 in two of three runs and reports at 42 of 50. It kept the unsalted MD5 passwords in every run; in its first it was the only agent to look something up on the web, the PHP manual page for fputcsv.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 638 8 of 13 flaws fixed no false positive contract kept 22/22 checks 29min US$ 2.43 Scorecard
- Run 2 · 663 8 of 13 flaws fixed no false positive contract kept 22/22 checks 24min US$ 2.88 Scorecard
- Run 3 · 607 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 26min US$ 2.57 Scorecard
-
10
Grok 4.6 xAI · effort high 3 of 3 runs (633 · 640 · 620) no published cutoff or earlier release633 of 1000 LEB Silver
- Security 194/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 25/100
- Compatibility 100/100
- Explanation 35/50
Comment and details
Nearly Grok 4.7's score for about an eighth of the cost (US$ 0.28 to 0.31 a run), and steady: 620 to 640, 9 flaws found in each run, no contract broken. It never touched MD5 or the CSV formula injection.
- Full marks
- Performance, Compatibility
- Zero
- Architecture
- Never fixed, in any run
- Formula injection in the CSV export; Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 633 7 of 13 flaws fixed no false positive contract kept 22/22 checks 6min US$ 0.31 Scorecard
- Run 2 · 640 8 of 13 flaws fixed no false positive contract kept 22/22 checks 5min US$ 0.31 Scorecard
- Run 3 · 620 7 of 13 flaws fixed no false positive contract kept 22/22 checks 6min US$ 0.28 Scorecard
-
11
Grok 4.7 xAI · effort xhigh 3 of 3 runs (640 · 617 · 631)631 of 1000 LEB Silver
- Security 185/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 25/100
- Compatibility 100/100
- Explanation 42/50
Comment and details
More effort scored slightly lower than Grok 4.7 at high, at a higher cost: it fixed 7 or 8 flaws a run and left MD5 and the hardcoded secrets in every one. Its reports, 42 to 44 of 50, are the best explanations from outside Anthropic and OpenAI.
- Full marks
- Performance, Compatibility
- Zero
- Architecture
- Never fixed, in any run
- Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 640 8 of 13 flaws fixed no false positive contract kept 22/22 checks 37min US$ 3.82 Scorecard
- Run 2 · 617 7 of 13 flaws fixed no false positive contract kept 22/22 checks 29min US$ 3.59 Scorecard
- Run 3 · 631 7 of 13 flaws fixed no false positive contract kept 22/22 checks 33min US$ 3.31 Scorecard
-
12
Kimi K3 Moonshot AI · effort high 3 of 3 runs (629 · 634 · 574) no published cutoff or earlier release629 of 1000 LEB Silver
- Security 198/250
- Architecture 0/200
- Bugs 150/150
- Performance 150/150
- Clean code 25/100
- Compatibility 70/100
- Explanation 36/50
Comment and details
The strongest Kimi by a wide margin, with bugs and performance at full marks in its official run, for US$ 0.41 to 1.02 a run. Every run changed one business value (−30), and none fixed the CSV formula injection.
- Full marks
- Bugs, Performance
- Zero
- Architecture
- Never fixed, in any run
- Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 629 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 7min US$ 0.41 Scorecard
- Run 2 · 634 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 21min US$ 1.02 Scorecard
- Run 3 · 574 7 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 15min US$ 0.81 Scorecard
-
13
GPT-6-astra OpenAI · effort xhigh 3 of 3 runs (661 · 596 · 628)628 of 1000 LEB Silver
- Security 228/250
- Architecture 0/200
- Bugs 139/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 41/50
Comment and details
Single-agent GPT-6-astra, 38 points below its ultra setting. Its runs disagree on what to fix: the first fixed the CSV formula injection and kept MD5, the official one migrated MD5 to password_hash and left the injection alone.
- Full marks
- Performance
- Zero
- Architecture, Clean code
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 661 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 17min Scorecard
- Run 2 · 596 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 12min Scorecard
- Run 3 · 628 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 14min Scorecard
-
14
GLM-5.3 Prime Z.AI · effort high 3 of 3 runs (635 · 628 · 541) no published cutoff or earlier release628 of 1000 LEB Silver
- Security 211/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 38/50
Comment and details
The strongest GLM, at about US$ 2 a run. Its first run fixed all four flaws the probes cover, the CSV formula injection among them; the third changed two business values and fell to 541.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 635 9 of 13 flaws fixed 1 false positive contract kept 22/22 checks 15min US$ 1.69 Scorecard
- Run 2 · 628 8 of 13 flaws fixed no false positive contract kept 22/22 checks 26min US$ 2.13 Scorecard
- Run 3 · 541 7 of 13 flaws fixed no false positive 2 business values changed 22/22 checks 17min US$ 1.77 Scorecard
-
15
GPT-5.6-terra OpenAI · effort xhigh 3 of 3 runs (625 · 611 · 645)625 of 1000 LEB Silver
- Security 211/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 35/50
Comment and details
Reports fewer flaws than it fixes: its first run listed 7 and fixed 9, the rest done in passing without a word in the report. Compatibility held in two of three runs, which keeps it 15th despite the short report.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
-
16
GLM-5.3-Flash Z.AI · effort high 1 of 3 runs · not official no published cutoff or earlier release624 of 1000 LEB Silver
- Security 211/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 34/50
Comment and details
One run so far: 624 for US$ 0.08, about a tenth of GLM-5.3's cost for a slightly higher score. It left the CSV formula injection and the hardcoded secrets in place.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 624 8 of 13 flaws fixed no false positive contract kept 22/22 checks 15min US$ 0.08 Scorecard
-
17
GLM-5.3 Z.AI · effort high 3 of 3 runs (629 · 621 · 604) no published cutoff or earlier release621 of 1000 LEB Silver
- Security 203/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 39/50
Comment and details
Steady (604 to 629) and cheap (US$ 0.47 to 0.73 a run), with compatibility at 100 in all three runs. It never fixed MD5 or the CSV formula injection.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 629 8 of 13 flaws fixed no false positive contract kept 22/22 checks 17min US$ 0.73 Scorecard
- Run 2 · 621 7 of 13 flaws fixed no false positive contract kept 22/22 checks 15min US$ 0.51 Scorecard
- Run 3 · 604 7 of 13 flaws fixed no false positive contract kept 22/22 checks 13min US$ 0.47 Scorecard
-
18
DeepSeek V4 Flash DeepSeek · effort xhigh 1 of 3 runs · not official no published cutoff or earlier release617 of 1000 LEB Silver
- Security 207/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 31/50
Comment and details
One run so far, through Novita: 617 for US$ 0.18, against 282 for the same model at high. It fixed 7 flaws and left the CSV injection, MD5 and the hardcoded secrets in place.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 617 7 of 13 flaws fixed no false positive contract kept 22/22 checks 28min US$ 0.18 Scorecard
-
19
GPT-6.1-sol OpenAI · effort ultra 3 of 3 runs (597 · 656 · 616)616 of 1000 LEB Silver
- Security 228/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 39/50
Comment and details
Ultra effort scored 45 points below GPT-6.1-sol at xhigh, across runs from 597 to 656. With three subagents its first run kept MD5; with two, its second fixed 10 flaws, the CSV formula injection and MD5 among them.
- Full marks
- Performance
- Zero
- Architecture, Clean code
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 597 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 24min Scorecard
- Run 2 · 656 10 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 13min Scorecard
- Run 3 · 616 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 11min Scorecard
-
20
GPT-5.6-sol OpenAI · effort xhigh 3 of 3 runs (612 · 612 · 608)612 of 1000 LEB Silver
- Security 246/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 32/50
Comment and details
Security at 246 of 250 and the CSV formula injection fixed, but the fix also turned the '-' of a ticket with no technician into "'-", a new bug, and every run changed one business value (−30). The most repeatable GPT: 608 to 612.
- Full marks
- Performance
- Zero
- Architecture, Clean code
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 612 10 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 7min Scorecard
- Run 2 · 612 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 15min Scorecard
- Run 3 · 608 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 15min Scorecard
-
21
DeepSeek V4.1 Flash DeepSeek · effort high 3 of 3 runs (625 · 612 · 597) no published cutoff or earlier release612 of 1000 LEB Silver
- Security 203/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 30/50
Comment and details
The cheapest agent here, US$ 0.01 to 0.04 a run, two of them under two minutes, for 597 to 625. Compatibility held in all three; the lowest explanation score among the agents above 600 (30 of 50).
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 625 9 of 13 flaws fixed no false positive contract kept 22/22 checks 9min US$ 0.04 Scorecard
- Run 2 · 612 8 of 13 flaws fixed no false positive contract kept 22/22 checks 1min US$ 0.01 Scorecard
- Run 3 · 597 7 of 13 flaws fixed no false positive contract kept 22/22 checks 1min US$ 0.02 Scorecard
-
22
GPT-5.6-luna OpenAI · effort xhigh 3 of 3 runs (599 · 601 · 624)601 of 1000 LEB Silver
- Security 220/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 32/50
Comment and details
The smallest GPT-5.6, within 24 points of terra and sol across 599 to 624. It changed one business value in each run (−30).
- Full marks
- Performance
- Zero
- Architecture, Clean code
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 599 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 12min Scorecard
- Run 2 · 601 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 10min Scorecard
- Run 3 · 624 10 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 14min Scorecard
-
23
GLM-5.3-FlashX Z.AI · effort high 1 of 3 runs · not official no published cutoff or earlier release597 of 1000 LEB Bronze
- Security 185/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 33/50
Comment and details
One run so far: 597 for US$ 0.10, 27 points below GLM-5.3-Flash. It fixed 7 flaws and left the CSV injection, MD5 and the hardcoded secrets in place.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 597 7 of 13 flaws fixed no false positive contract kept 22/22 checks 9min US$ 0.10 Scorecard
-
24
Gemini 3.8 Flash Google · effort high 3 of 3 runs (687 · 588 · 100) no published cutoff or earlier release588 of 1000 LEB Bronze
- Security 181/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 28/50
Comment and details
Three very different runs: 687 (it would have placed 6th), 588, and 100 from a run that stopped after ten minutes on a database server it started itself. The two complete runs kept MD5 and both secrets and left the CSV export serving every client.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority
Run by run
- Run 1 · 687 8 of 13 flaws fixed no false positive contract kept 22/22 checks 13min US$ 1.00 Scorecard
- Run 2 · 588 7 of 13 flaws fixed no false positive contract kept 22/22 checks 25min US$ 0.85 Scorecard
- Run 3 · 100 0 of 13 flaws fixed no false positive contract kept 22/22 checks 10min US$ 0.28 Scorecard
-
25
Gemini 3.7 Flash Google · effort high 3 of 3 runs (587 · 587 · 556) no published cutoff or earlier release587 of 1000 LEB Bronze
- Security 181/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 27/50
Comment and details
The generation before Gemini 3.8 Flash, one point behind it and far steadier (556 to 587). Its first run is the only one outside Claude to flatten the nested ifs; its second searched the VM for the answer key, which is not there.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 587 7 of 13 flaws fixed no false positive contract kept 22/22 checks 12min US$ 0.72 Scorecard
- Run 2 · 587 7 of 13 flaws fixed no false positive contract kept 22/22 checks 16min US$ 0.60 Scorecard
- Run 3 · 556 7 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 7min US$ 0.49 Scorecard
-
26
GPT-5.5 OpenAI · effort xhigh 3 of 3 runs (558 · 568 · 536)558 of 1000 LEB Bronze
- Security 177/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 32/50
Comment and details
The previous GPT generation, 43 to 108 points below the GPT-5.6 and GPT-6 models: 8 flaws found and 7 fixed in every run, MD5 and the CSV formula injection always left in place.
- Full marks
- Performance
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
-
27
Gemini 3.8 Flash Google · effort medium 1 of 3 runs · not official no published cutoff or earlier release550 of 1000 LEB Bronze
- Security 147/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 24/50
Comment and details
One run so far, in 4min 34s for US$ 0.19: 550. It fixed 6 flaws and claimed SQL injection in two functions that only take integers.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 550 6 of 13 flaws fixed 2 false positives contract kept 22/22 checks 5min US$ 0.19 Scorecard
-
28
Qwen3 Coder Next Alibaba (Qwen team) · default effort (not configurable) 1 of 3 runs · not official507 of 1000 LEB Bronze
- Security 125/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 18/50
Comment and details
The weakest explanation (18 of 50) and the worst calibration here: three flaws that cannot exist reported at confidence 100, and no tests run. It is, on the other hand, the only one of the bottom nine that fixed the N+1 query.
- Full marks
- Performance, Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Reflected XSS in the search; Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; Any ticket readable by id (IDOR); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 507 5 of 13 flaws fixed 3 false positives contract kept 22/22 checks 16min US$ 1.53 Scorecard
-
29
DeepSeek V4 Pro DeepSeek · effort high 3 of 3 runs (604 · 496 · 432) no published cutoff or earlier release496 of 1000 LEB Bronze
- Security 181/250
- Architecture 0/200
- Bugs 129/150
- Performance 56/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 30/50
Comment and details
The larger DeepSeek: all three runs (604, 496 and 432) score below DeepSeek V4.1 Flash's official 612. The first rated SQL injection at confidence 100 in two functions that only take integers; the other two left the N+1 query in place.
- Full marks
- Compatibility
- Zero
- Architecture, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 604 8 of 13 flaws fixed 2 false positives contract kept 22/22 checks 8min US$ 0.03 Scorecard
- Run 2 · 496 6 of 13 flaws fixed no false positive contract kept 22/22 checks 23min US$ 0.42 Scorecard
- Run 3 · 432 6 of 13 flaws fixed 1 false positive contract kept 22/22 checks 12min US$ 0.04 Scorecard
-
30
MiniMax-M3 MiniMax · effort thinking 3 of 3 runs (462 · 616 · 434)462 of 1000 LEB Bronze
- Security 220/250
- Architecture 0/200
- Bugs 150/150
- Performance 0/150
- Clean code 25/100
- Compatibility 40/100
- Explanation 27/50
Comment and details
Fixed 8 flaws in every run, with bugs at full marks, but changed business values (compatibility at 40 in its official run) and left the N+1 query. Its second run scored 616; its third broke four of the 22 characterization checks.
- Full marks
- Bugs
- Zero
- Architecture, Performance
- Never fixed, in any run
- Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority
Run by run
- Run 1 · 462 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 4min US$ 0.12 Scorecard
- Run 2 · 616 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 6min US$ 0.10 Scorecard
- Run 3 · 434 8 of 13 flaws fixed 1 false positive 1 business value changed 18/22 checks 5min US$ 0.13 Scorecard
-
31
Kimi K2.7 Code Moonshot AI · default effort (not configurable) 3 of 3 runs (415 · 415 · 512)415 of 1000 LEB Bronze
- Security 203/250
- Architecture 0/200
- Bugs 86/150
- Performance 0/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 26/50
Comment and details
Left the N+1 query in every run and reported flaws that do not exist in all three, SQL injection in functions that only take integers among them. Its third run, 512, fixed one flaw more than the other two.
- Full marks
- Compatibility
- Zero
- Architecture, Performance, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Unsalted MD5 passwords; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 415 6 of 13 flaws fixed 2 false positives contract kept 22/22 checks 14min US$ 0.48 Scorecard
- Run 2 · 415 6 of 13 flaws fixed 1 false positive contract kept 22/22 checks 4min US$ 0.25 Scorecard
- Run 3 · 512 7 of 13 flaws fixed 2 false positives contract kept 22/22 checks 15min US$ 0.81 Scorecard
-
32
GPT-5.3-Codex OpenAI · effort xhigh 1 of 3 runs · not official403 of 1000 LEB Bronze
- Security 147/250
- Architecture 0/200
- Bugs 129/150
- Performance 0/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 27/50
Comment and details
One run so far, through opencode: 403, the lowest GPT. It found 6 flaws and fixed 5, leaving the N+1 query, the session fixation at login and MD5 in place, with no false positive.
- Full marks
- Compatibility
- Zero
- Architecture, Performance, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; Secrets hardcoded in the config; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 403 5 of 13 flaws fixed no false positive contract kept 22/22 checks 7min US$ 0.54 Scorecard
-
33
Kimi K2.7 Code (highspeed) Moonshot AI · default effort (not configurable) 3 of 3 runs (402 · 363 · 402) no published cutoff or earlier release402 of 1000 LEB Bronze
- Security 155/250
- Architecture 0/200
- Bugs 129/150
- Performance 0/150
- Clean code 25/100
- Compatibility 70/100
- Explanation 23/50
Comment and details
The fast variant of Kimi K2.7 Code, 13 points below it: 5 flaws fixed in every run, two flaws that do not exist reported in each, and the N+1 query always left in place.
- Zero
- Architecture, Performance
- Never fixed, in any run
- Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; Secrets hardcoded in the config; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 402 5 of 13 flaws fixed 2 false positives 1 business value changed 22/22 checks 2min US$ 0.38 Scorecard
- Run 2 · 363 5 of 13 flaws fixed 2 false positives 1 business value changed 22/22 checks 7min US$ 0.68 Scorecard
- Run 3 · 402 5 of 13 flaws fixed 2 false positives contract kept 22/22 checks 10min US$ 1.16 Scorecard
-
34
GLM-5.2 Z.AI · effort high 1 of 3 runs · not official388 of 1000 Failed
- Security 172/250
- Architecture 0/200
- Bugs 86/150
- Performance 0/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 30/50
Comment and details
One run so far: 388, below the pass line and 233 points under GLM-5.3. It left the N+1 query, the leaked file handle and the session fixation in place; what it did report was accurate, with no false positive.
- Full marks
- Compatibility
- Zero
- Architecture, Performance, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; File handle leaked on the error path; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 388 5 of 13 flaws fixed no false positive contract kept 22/22 checks 6min US$ 0.35 Scorecard
-
35
Nex N2.5 Pro Nex AGI · effort high 2 of 3 runs (317 · 428) · not official no published cutoff or earlier release317 of 1000 Failed
- Security 194/250
- Architecture 0/200
- Bugs 107/150
- Performance 75/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 31/50
Comment and details
Too slow to finish three runs: 10h 03min and 19h 35min, and a third voided after more than 21 hours when the operator's connection failed, so it stays at two and publishes the lower. Both runs were strong on security and both broke the contract: the first made every function return nothing without a logged-in user (8 checks, −160, 317), the second made the CSV export throw for a caller that has already printed output (3 checks, 428).
- Zero
- Architecture, Clean code
- Never fixed, in any run
- A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
-
36
Claude Haiku 4.5 Anthropic · default effort (not configurable) 3 of 3 runs (317 · 369 · 232)317 of 1000 Failed
- Security 138/250
- Architecture 0/200
- Bugs 86/150
- Performance 0/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 23/50
Comment and details
The small Claude, at US$ 0.21 to 0.40 a run, below the pass line in all three. Its first run's visibility fix hides a client's own tickets, comparing an integer with the string mysqli returns; its third broke three characterization checks.
- Zero
- Architecture, Performance, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; File handle leaked on the error path; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
- Run 1 · 317 4 of 13 flaws fixed 2 false positives 1 business value changed 22/22 checks 3min US$ 0.21 Scorecard
- Run 2 · 369 5 of 13 flaws fixed 2 false positives contract kept 22/22 checks 7min US$ 0.40 Scorecard
- Run 3 · 232 5 of 13 flaws fixed 1 false positive 1 business value changed 19/22 checks 10min US$ 0.27 Scorecard
-
37
DeepSeek V4 Flash DeepSeek · effort high 2 of 3 runs (612 · 282) · not official no published cutoff or earlier release282 of 1000 Failed
- Security 99/250
- Architecture 0/200
- Bugs 96/150
- Performance 0/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 22/50
Comment and details
Two runs that could hardly differ more: the first fixed 9 flaws for 612, the second's SQL-injection fix throws on every search and its client CSV comes out on one line, 282. It publishes the lower; which weights the hosts served under this name is not certain.
- Full marks
- Compatibility
- Zero
- Architecture, Performance, Clean code
- Never fixed, in any run
- Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
Run by run
Flaw by flaw
What each agent found and fixed among the planted flaws. The hard ones are flaws of absence — a missing authorization check, a session never regenerated, a file left open on the error path.
The table shows the top 10 of 37 agents. Every agent's result, flaw by flaw, is in its scorecard.
- fixed
- found, not fixed
- missed
| Flaw | Claude Sonnet 5.5 · xhigh | Claude Sonnet 5.5 · max | Claude Sonnet 5.5 · max (ultracode) | Claude Fable 5.1 | Claude Opus 5.5 | GPT-6-astra · ultra | GPT-6.1-sol · xhigh | GPT-6.1-sol pro | Grok 4.7 · high | Grok 4.6 |
|---|---|---|---|---|---|---|---|---|---|---|
| SQL injection in the search | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed |
| Reflected XSS in the search | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed |
| Formula injection in the CSV export | 6/6 fixed | 6/6 fixed | 6/6 fixed | 2/6 found, not fixed | 2/6 found, not fixed | 6/6 fixed | 6/6 fixed | 2/6 found, not fixed | 6/6 fixed | 0/6 missed |
| Session fixation at login | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed |
| Unsalted MD5 passwords | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 3/8 found, not fixed | 8/8 fixed | 8/8 fixed | 3/8 found, not fixed | 3/8 found, not fixed |
| Secrets hardcoded in the config | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 3/8 found, not fixed | 6/8 found, not fixed |
| Any ticket readable by id (IDOR) | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 9/10 fixed | 9/10 fixed | 9/10 fixed | 10/10 fixed | 10/10 fixed |
| Division by zero in the SLA average | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed |
| File handle leaked on the error path | 5/6 fixed | 6/6 fixed | 4/6 fixed | 5/6 fixed | 5/6 fixed | 4/6 fixed | 4/6 fixed | 5/6 fixed | 5/6 fixed | 4/6 fixed |
| One query per ticket for the technician (N+1) | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed |
| A dispatcher that does everything | 2/10 found, not fixed | 1/10 found, not fixed | 0/10 missed | 0/10 missed | 4/10 found, not fixed | 2/10 found, not fixed | 2/10 found, not fixed | 0/10 missed | 0/10 missed | 0/10 missed |
| Magic numbers for status and priority | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed |
| Four levels of nested ifs | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 0/8 missed | 2/8 found, not fixed | 0/8 missed | 2/8 found, not fixed | 0/8 missed | 2/8 found, not fixed |
What stood out
- Six agents fixed the formula injection in the CSV (SEC-008) — Sonnet 5.5 in all three of its settings, GPT-6.1-sol, GPT-5.6-sol and Grok 4.7 — with the fix the answer key expects; Sonnet 5.5 is the only model with security at 250 of 250, in all three of its settings. GPT-5.6-sol also turned the
-of a ticket with no technician into'-, a new bug (−15). - Architecture was the weakest category for everyone: 50 of 200 for Opus 5.5, 25 for Sonnet 5.5 at xhigh and for GPT-6.1-sol, 12 for Sonnet 5.5 at max, and 0 for the other twenty-eight. Nobody split the dispatcher that does everything; those three named it and declined to restructure it.
- Two published runs broke the contract mechanically. All thirty-seven stayed on mysqli and reported no decoy, and thirty-five kept the 22 characterization checks green; DeepSeek V4 Flash publishes a run whose search throws, and Nex N2.5 Pro one whose functions return nothing without a logged-in user (8 checks broken, −160). Judgement is what separated the rest: twenty-four kept compatibility at 100, while the other thirteen each changed at least one business value (−30 each), most often the SLA average, scoped to each client.
- Twenty-six scores are official: Claude Sonnet 5.5 at xhigh, 809 (runs of 825, 809 and 724), Claude Sonnet 5.5 at max, 807 (807, 773 and 820), Claude Sonnet 5.5 at max in multi-agent mode, 774 (774, 820 and 759), Claude Opus 5.5, 717 (711, 717 and 805), GPT-6.1-sol, 661 (666, 653 and 661), GPT-6.1-sol at ultra, 616 (597, 656 and 616), Grok 4.7, 638 (638, 663 and 607), Grok 4.6, 633 (633, 640 and 620), Grok 4.7 at xhigh, 631 (640, 617 and 631), GPT-6-astra at ultra, 666 (668, 666 and 651), Kimi K3 at high, 629 (629, 634 and 574), GPT-6-astra, 628 (661, 596 and 628), GPT-5.6-terra, 625 (625, 611 and 645), GLM-5.3 Prime, 628 (635, 628 and 541), GLM-5.3, 621 (629, 621 and 604), DeepSeek V4.1 Flash, 612 (625, 612 and 597), GPT-5.6-sol, 612 (612, 612 and 608), GPT-5.6-luna, 601 (599, 601 and 624), Gemini 3.8 Flash at high, 588 (687, 588 and 100), GPT-5.5, 558 (558, 568 and 536), DeepSeek V4 Pro, 496 (604, 496 and 432), MiniMax-M3 with the thinking variant, 462 (462, 616 and 434), Kimi K2.7 Code, 415 (415, 415 and 512), Kimi K2.7 Code highspeed, 402 (402, 363 and 402), Gemini 3.7 Flash at high, 587 (587, 587 and 556), and Claude Haiku 4.5, 317 (317, 369 and 232), each the median of three runs. A single run can sit more than 100 points from the median: Opus's third scored 805, Sonnet's third 724, DeepSeek V4 Pro's first 604, Gemini 3.8 Flash's third, which stopped after ten minutes, 100, and astra's first, 661, had placed it 5th, from judgement calls such as which flaws to leave unfixed. Claude Fable 5.1 has two runs, 781 and 764, and publishes the lower, 4th
- Gemini 3.8 Flash at high is official at 588, 24th, Bronze, the median of three very different runs (687, 588 and 100). Gemini 3.7 Flash, the generation before, is official at 587 (587, 587 and 556), 25th. Its first run, 6th, fixed 8 of the 13 planted flaws and is the only one outside Claude to flatten the nested ifs (CLN-007); its second fixed 7, and searched the VM for the answer key, by name and by the hash quoted in the task. The key is not on the VM, and the search turned up nothing the run did not already have. Its third stopped after ten minutes, on a database server it started that kept its own shell command from ever returning: no report, the code untouched, scored as delivered. The two complete runs kept MD5 and both secrets and left the CSV export serving every client. At medium effort the same model scored 550 (27th) in 4min 34s: it fixed 6 flaws and claimed SQL injection in two functions that only take integers.
- GPT-6-astra at ultra is the strongest GPT agent here, official at 666 (runs of 668, 666 and 651), 6th, with two or three subagents in 10 to 15 minutes. GPT-6.1-sol follows, official at 661 (666, 653 and 661), 7th, and GPT-6.1-sol pro at 654, 8th, from one run. GPT-6-astra at xhigh, official at 628, is 13th. Its first run fixed the CSV injection and kept MD5; its official run did the opposite, migrating MD5 to
password_hashand leaving the CSV injection alone; GPT-6.1-sol's official run fixed both - Multi-agent work matched one agent at best, at many times the cost. Claude Sonnet 5.5 in Claude Code's multi-agent mode, at max effort, scored 774 in its first run (7 workflows, 68 subagents, 7h 55min, US$ 233), 820 in its second (2 workflows, 105 subagents, 2h, US$ 171) and 759 in its third (3 workflows, 145 subagents, 3h 34min, US$ 119). It is official at 774, 3rd, Gold. The same model at the same max effort as a single agent took about half an hour a run and is official at 807 (runs of 807, 773 and 820), 2nd; at xhigh it is official at 809. Its third report scored 47 of 50, the best explanation here, and no run touched the architecture. In each run a few subagents (nine, eight, then seven) fell back to an older Sonnet after a safety classifier stopped them; none of any delivery came from them.
- GPT-6.1-sol at ultra is official 45 points below itself at xhigh (616 against 661), the median of three runs spread over 59 points (597, 656 and 616). With three subagents, its first found 9 of the 13 planted flaws, kept MD5 and did not name the dispatcher; with two, its second fixed 10, the CSV injection and MD5 among them. Each took under half an hour, not the 8 hours of the multi-agent Sonnet.
- Grok 4.7 is the strongest model here from outside Anthropic and OpenAI (638, 9th, Silver, official over three runs of 638, 663 and 607), with an explanation at the level of the best GPT reports (42 of 50). In its first run it is also the only agent that used the web, to read the PHP manual page for fputcsv; the VM blocks GitHub, not the web. At xhigh, Grok 4.7 is official at 631 (640, 617 and 631), 11th, below its own high effort, with reports scored 42 to 44 of 50, the best explanations from outside Anthropic and OpenAI. Grok 4.6 follows, official at 633 (10th; runs of 633, 640 and 620), for about an eighth of the cost (US$ 0.31 against 2.43).
- GLM-5.3 Prime is the strongest GLM model (628, 14th, Silver), official over three runs through two hosts (635, 628 and 541): the first fixed all four probe-covered flaws, the CSV injection among them, for US$ 1.69; the second left the injection alone. GLM-5.3 is official at 621 (17th), the median of three runs (629, 621 and 604), the last two served by another host; GLM-5.3-Flash scores 624 for about a tenth of the cost, and GLM-5.3-FlashX 597
- DeepSeek V4.1 Flash is the cheapest agent here (US$ 0.01 to 0.04 a run; its last two took under two minutes each) and is official at 612 (runs of 625, 612 and 597), 21st. DeepSeek V4 Flash at xhigh, through Novita, scored 617 in its first run (18th). At high, run through two other hosts, it publishes 282, the lower of its two runs (612 and 282), last: the first fixed eight flaws fully, while the second's SQL-injection fix throws on every search and its client CSV comes out on one line. DeepSeek now routes the V4 Flash name to V4.1 on its own API, so which weights those hosts served is not certain. DeepSeek V4 Pro, the larger model, is official at 496, the median of three runs through two hosts (604, 496 and 432), all below DeepSeek V4.1 Flash's official 612: the first rates SQL injection at confidence 100 in two functions that only take integers, and the other two left the N+1 query in place.
- Twenty-four agents did not run at xhigh. Kimi K2.7 Code (in both of Moonshot's ids), Qwen3 Coder Next and Claude Haiku 4.5 have no effort setting and ran at their model's default; the two Grok models, the five GLM models, the three DeepSeek models and Gemini 3.8 and 3.7 Flash and Nex N2.5 Pro ran at high, in opencode, and Gemini 3.8 Flash also at medium; MiniMax-M3 ran in opencode with its thinking variant and Kimi K3 with its high variant; Sonnet 5.5 ran twice at max, once in multi-agent mode, and GPT-6.1-sol and GPT-6-astra once each at ultra. The rest ran at xhigh.
- Qwen3 Coder Next, DeepSeek V4 Pro, MiniMax-M3 with the thinking variant, Kimi K2.7 Code, GPT-5.3-Codex, Kimi K2.7 Code highspeed, GLM-5.2, Nex N2.5 Pro, Claude Haiku 4.5 and DeepSeek V4 Flash close the table (507, 496, 462, 415, 403, 402, 388, 317, 317 and 282; the last four are below the pass line). Nex N2.5 Pro is too slow to finish three runs: 10h 03min and 19h 35min, the longest single-agent runs here, and a third voided after more than 21 hours, so it stays at two; both fixed most of the security flaws, both broke the contract, and it publishes the lower, 317 (317 and 428). Claude Haiku 4.5 is official at 317 (runs of 317, 369 and 232); in the first, 2.5 minutes and US$ 0.21, its visibility fix hides a client's own tickets on the main page, comparing an integer with the string mysqli returns. All but Qwen3 Coder Next left the N+1 query in place. Qwen3 Coder Next has the weakest explanation (18 of 50) and the worst calibration (Brier 0.301): it reported three flaws that cannot exist at confidence 100, SQL injection in two functions that only take integers among them, as Kimi K2.7 Code did. It also ran no tests.
- Places 9 to 23 sit within 41 points (638 to 597), inside the noise of a single run. GPT-5.6-terra reported the fewest planted flaws and still ranks 15th, on compatibility and on what it did fix.
- The GPT models are among the best calibrated (Brier 0.015 or less): fewer findings, each stated with high confidence and nearly all real.
Finding is not fixing
Side by side, the three Claude models find nearly the same flaws and fix different numbers of them. One total hides which of the two it is measuring.
- They find nearly the same. In the run that counts for each, Sonnet 5.5 at xhigh reported 12 of the 13 planted flaws and Fable 5.1 and Opus 5.5 11; across their eight runs, each found 11 or 12
- Sonnet 5.5 fixes more, in the run that counts. Its official run fixed 11 of the 13, against 9 for Opus's official run and 10 for Fable's. Run by run it varies: Sonnet fixed 11, 11 and 8, Opus 9 in each of its three, Fable 9 and 10. Fable and Opus found and explained the formula injection in the CSV and left it in place, to protect the file's consumers; Sonnet fixed it in two of its three runs, prefixing only the cells that start a formula. Its 92-point lead over Opus's official 717 comes from clean code (+100), not from security (+17); its 45-point lead over Fable's 764 is security (+17) and architecture (+25)
- More compute did not change the pattern. The single-agent Claude runs took 16 to 27 minutes each, and the three at the same max effort scored 807, 773 and 820. Sonnet 5.5 in multi-agent mode took 7h 55min, 2h and 3h 34min and US$ 233, 171 and 119 for 774, 820 and 759, and is official at 774 against the single agent's 807: at best it tied the single-agent 820, with the same fixes and a deeper check of its own code, and in none of its three runs did the dispatcher reach its report.
- How to read it. On this task the three Claude models see the same problems; what separates them is how far they go in fixing them within the contract. In the median, Sonnet went further; a single run can reverse it, as Sonnet's third did. One task, two or three runs per model: a pattern worth testing, not a verdict