Skip to content

The next level, LEB-300, is an exploratory pilot: its first aggregate result is published.

AI Benchmark · LEB

LEB-100 the reference instance

The first level of LEB: a legacy PHP application of about 300 lines, solved by every agent below from the same package. How a run works, how it is scored and how this level differs from LEB-300 is on the Benchmark page.

Results

Every delivery in a table solved the same package, byte for byte — the same SHA-256 — so the numbers compare like for like.

LEB-100-A v1.1 · Support-ticket panel of an internet provider

The support-ticket panel of an internet provider, written in 2013-style PHP: data-access functions and an index.php that routes, authorizes and builds the HTML. About 300 lines on PHP 8, mysqli and MySQL 8, with 13 planted flaws and 2 decoys.

mode A · 30 turns edition 2026 evaluated 2026-10-06 matrix 68088abdb7bc…
  1. 1
    Claude Sonnet 5.5 Anthropic · effort xhigh 3 of 3 runs (825 · 809 · 724)
    809 of 1000 LEB Gold
    • Security 250/250
    • Architecture 25/200
    • Bugs 139/150
    • Performance 150/150
    • Clean code 100/100
    • Compatibility 100/100
    • Explanation 45/50

    Penalties: none Discovery 91.7 Brier 0.033 Cost US$ 3.16 a run Scorecard

    Comment and details

    The most complete single-agent result: security at 250 of 250 in its first two runs, performance at full marks, compatibility untouched, and the CSV formula injection fixed the way the answer key expects. Its third run left three security flaws found but unfixed and fell to 724; like every agent, it left the dispatcher whole.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Security, Performance, Clean code, Compatibility
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority
    Run by run
    • Run 1 · 825 11 of 13 flaws fixed no false positive contract kept 22/22 checks 19min US$ 3.60 Scorecard
    • Run 2 · 809 11 of 13 flaws fixed no false positive contract kept 22/22 checks 23min US$ 2.96 Scorecard
    • Run 3 · 724 8 of 13 flaws fixed no false positive contract kept 22/22 checks 23min US$ 2.91 Scorecard
  2. 2
    Claude Sonnet 5.5 Anthropic · effort max 3 of 3 runs (807 · 773 · 820)
    807 of 1000 LEB Gold
    • Security 250/250
    • Architecture 12/200
    • Bugs 150/150
    • Performance 150/150
    • Clean code 100/100
    • Compatibility 100/100
    • Explanation 45/50

    Penalties: none Discovery 91.7 Brier 0.035 Cost US$ 5.09 a run Scorecard

    Comment and details

    The most consistent of the top three: 773 to 820 across three runs, bugs and performance at full marks in all of them, no false positive and no contract broken. More effort than xhigh bought nothing measurable, at about one and a half times the cost.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Security, Bugs, Performance, Clean code, Compatibility
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority
    Run by run
    • Run 1 · 807 11 of 13 flaws fixed no false positive contract kept 22/22 checks 27min US$ 3.86 Scorecard
    • Run 2 · 773 10 of 13 flaws fixed no false positive contract kept 22/22 checks 29min US$ 4.75 Scorecard
    • Run 3 · 820 11 of 13 flaws fixed no false positive contract kept 22/22 checks 41min US$ 6.65 Scorecard
  3. 3
    Claude Sonnet 5.5 Anthropic · effort max (ultracode) 3 of 3 runs (774 · 820 · 759)
    774 of 1000 LEB Gold
    • Security 250/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 100/100
    • Compatibility 100/100
    • Explanation 45/50

    Penalties: none Discovery 87.5 Brier 0.008 Cost US$ 174.44 a run Scorecard

    Comment and details

    Claude Code's multi-agent mode: 68 to 145 subagents and 2h to 7h 55min a run, at US$ 119 to 233 each, for 774, 820 and 759. At best it tied the same model working alone in half an hour, with the best-calibrated report of the Claude agents and the same untouched architecture.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Security, Performance, Clean code, Compatibility
    Zero
    Architecture
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority
    Run by run
    • Run 1 · 774 11 of 13 flaws fixed no false positive contract kept 22/22 checks 7h 55min US$ 233.08 Scorecard
    • Run 2 · 820 11 of 13 flaws fixed no false positive contract kept 22/22 checks 2h US$ 171.18 Scorecard
    • Run 3 · 759 10 of 13 flaws fixed no false positive contract kept 22/22 checks 3h 34min US$ 119.07 Scorecard
  4. 4
    Claude Fable 5.1 Anthropic · effort xhigh 2 of 3 runs (781 · 764) · not official
    764 of 1000 LEB Gold
    • Security 233/250
    • Architecture 0/200
    • Bugs 139/150
    • Performance 150/150
    • Clean code 100/100
    • Compatibility 100/100
    • Explanation 42/50

    Penalties: none Discovery 87.5 Brier 0.021 Cost US$ 8.04 a run Scorecard

    Comment and details

    Finds as much as Sonnet (11 to 12 of the 13 flaws) and fixes less of it: it explained the CSV formula injection and left it in place on purpose, to protect the file's consumers. Two runs so far; its client switched part of a third attempt to another model, which voided it.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Clean code, Compatibility
    Zero
    Architecture
    Never fixed, in any run
    Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority
    Run by run
    • Run 1 · 781 9 of 13 flaws fixed no false positive contract kept 22/22 checks 16min Scorecard
    • Run 2 · 764 10 of 13 flaws fixed no false positive contract kept 22/22 checks 22min US$ 8.04 Scorecard
  5. 5
    Claude Opus 5.5 Anthropic · effort xhigh 3 of 3 runs (711 · 717 · 805)
    717 of 1000 LEB Silver
    • Security 233/250
    • Architecture 50/200
    • Bugs 139/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 45/50

    Penalties: none Discovery 87.5 Brier 0.024 Cost US$ 2.86 a run Scorecard

    Comment and details

    The highest architecture score in an official run (50 of 200), and the widest spread among the Claude models: 711 and 717, then 805. It fixed 9 flaws in every run and never flattened the nested ifs in the run that counts, which is where its gap to Sonnet comes from.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Clean code
    Never fixed, in any run
    Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority
    Run by run
    • Run 1 · 711 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 16min Scorecard
    • Run 2 · 717 9 of 13 flaws fixed no false positive contract kept 22/22 checks 18min US$ 3.01 Scorecard
    • Run 3 · 805 9 of 13 flaws fixed no false positive contract kept 22/22 checks 15min US$ 2.71 Scorecard
  6. 6
    GPT-6-astra OpenAI · effort ultra 3 of 3 runs (668 · 666 · 651)
    666 of 1000 LEB Silver
    • Security 224/250
    • Architecture 25/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 25/100
    • Compatibility 70/100
    • Explanation 43/50

    Penalties: none Discovery 79.2 Brier 0.000 Scorecard

    Comment and details

    The strongest GPT agent, with two or three subagents in 10 to 15 minutes a run and three runs within 17 points of each other. It changed one business value in every run (−30 each), and its reports are among the best calibrated here (Brier 0.000).

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 668 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 11min Scorecard
    • Run 2 · 666 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 15min Scorecard
    • Run 3 · 651 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 10min Scorecard
  7. 7
    GPT-6.1-sol OpenAI · effort xhigh 3 of 3 runs (666 · 653 · 661)
    661 of 1000 LEB Silver
    • Security 246/250
    • Architecture 25/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 41/50

    Penalties: none Discovery 75.0 Brier 0.000 Scorecard

    Comment and details

    Security at 246 of 250, the best outside Claude, with the CSV formula injection and MD5 both fixed in the official run. What holds it below the Claude models is one business value changed in each run (−30) and the nested ifs left as they were.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance
    Zero
    Clean code
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 666 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 27min Scorecard
    • Run 2 · 653 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 26min Scorecard
    • Run 3 · 661 10 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 24min Scorecard
  8. 8
    GPT-6.1-sol pro OpenAI · effort xhigh 1 of 3 runs · not official no published cutoff or earlier release
    654 of 1000 LEB Silver
    • Security 228/250
    • Architecture 0/200
    • Bugs 139/150
    • Performance 150/150
    • Clean code 25/100
    • Compatibility 70/100
    • Explanation 42/50

    Penalties: none Discovery 87.5 Brier 0.000 Cost US$ 1.01 a run Scorecard

    Comment and details

    One run so far, through opencode, at US$ 1.01: 654, close to GPT-6.1-sol in Codex CLI. It left the CSV formula injection in place and changed one business value (−30).

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance
    Zero
    Architecture
    Never fixed, in any run
    Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 654 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 16min US$ 1.01 Scorecard
  9. 9
    Grok 4.7 xAI · effort high 3 of 3 runs (638 · 663 · 607)
    638 of 1000 LEB Silver
    • Security 207/250
    • Architecture 0/200
    • Bugs 139/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 42/50

    Penalties: none Discovery 83.3 Brier 0.005 Cost US$ 2.63 a run Scorecard

    Comment and details

    The strongest model from outside Anthropic and OpenAI, with compatibility at 100 in two of three runs and reports at 42 of 50. It kept the unsalted MD5 passwords in every run; in its first it was the only agent to look something up on the web, the PHP manual page for fputcsv.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 638 8 of 13 flaws fixed no false positive contract kept 22/22 checks 29min US$ 2.43 Scorecard
    • Run 2 · 663 8 of 13 flaws fixed no false positive contract kept 22/22 checks 24min US$ 2.88 Scorecard
    • Run 3 · 607 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 26min US$ 2.57 Scorecard
  10. 10
    Grok 4.6 xAI · effort high 3 of 3 runs (633 · 640 · 620) no published cutoff or earlier release
    633 of 1000 LEB Silver
    • Security 194/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 25/100
    • Compatibility 100/100
    • Explanation 35/50

    Penalties: none Discovery 62.5 Brier 0.007 Cost US$ 0.30 a run Scorecard

    Comment and details

    Nearly Grok 4.7's score for about an eighth of the cost (US$ 0.28 to 0.31 a run), and steady: 620 to 640, 9 flaws found in each run, no contract broken. It never touched MD5 or the CSV formula injection.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture
    Never fixed, in any run
    Formula injection in the CSV export; Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 633 7 of 13 flaws fixed no false positive contract kept 22/22 checks 6min US$ 0.31 Scorecard
    • Run 2 · 640 8 of 13 flaws fixed no false positive contract kept 22/22 checks 5min US$ 0.31 Scorecard
    • Run 3 · 620 7 of 13 flaws fixed no false positive contract kept 22/22 checks 6min US$ 0.28 Scorecard
  11. 11
    Grok 4.7 xAI · effort xhigh 3 of 3 runs (640 · 617 · 631)
    631 of 1000 LEB Silver
    • Security 185/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 25/100
    • Compatibility 100/100
    • Explanation 42/50

    Penalties: none Discovery 75.0 Brier 0.001 Cost US$ 3.57 a run Scorecard

    Comment and details

    More effort scored slightly lower than Grok 4.7 at high, at a higher cost: it fixed 7 or 8 flaws a run and left MD5 and the hardcoded secrets in every one. Its reports, 42 to 44 of 50, are the best explanations from outside Anthropic and OpenAI.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture
    Never fixed, in any run
    Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 640 8 of 13 flaws fixed no false positive contract kept 22/22 checks 37min US$ 3.82 Scorecard
    • Run 2 · 617 7 of 13 flaws fixed no false positive contract kept 22/22 checks 29min US$ 3.59 Scorecard
    • Run 3 · 631 7 of 13 flaws fixed no false positive contract kept 22/22 checks 33min US$ 3.31 Scorecard
  12. 12
    Kimi K3 Moonshot AI · effort high 3 of 3 runs (629 · 634 · 574) no published cutoff or earlier release
    629 of 1000 LEB Silver
    • Security 198/250
    • Architecture 0/200
    • Bugs 150/150
    • Performance 150/150
    • Clean code 25/100
    • Compatibility 70/100
    • Explanation 36/50

    Penalties: none Discovery 75.0 Brier 0.008 Cost US$ 0.75 a run Scorecard

    Comment and details

    The strongest Kimi by a wide margin, with bugs and performance at full marks in its official run, for US$ 0.41 to 1.02 a run. Every run changed one business value (−30), and none fixed the CSV formula injection.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Bugs, Performance
    Zero
    Architecture
    Never fixed, in any run
    Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 629 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 7min US$ 0.41 Scorecard
    • Run 2 · 634 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 21min US$ 1.02 Scorecard
    • Run 3 · 574 7 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 15min US$ 0.81 Scorecard
  13. 13
    GPT-6-astra OpenAI · effort xhigh 3 of 3 runs (661 · 596 · 628)
    628 of 1000 LEB Silver
    • Security 228/250
    • Architecture 0/200
    • Bugs 139/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 41/50

    Penalties: none Discovery 83.3 Brier 0.002 Scorecard

    Comment and details

    Single-agent GPT-6-astra, 38 points below its ultra setting. Its runs disagree on what to fix: the first fixed the CSV formula injection and kept MD5, the official one migrated MD5 to password_hash and left the injection alone.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance
    Zero
    Architecture, Clean code
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 661 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 17min Scorecard
    • Run 2 · 596 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 12min Scorecard
    • Run 3 · 628 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 14min Scorecard
  14. 14
    GLM-5.3 Prime Z.AI · effort high 3 of 3 runs (635 · 628 · 541) no published cutoff or earlier release
    628 of 1000 LEB Silver
    • Security 211/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 38/50

    Penalties: none Discovery 70.8 Brier 0.029 Cost US$ 1.86 a run Scorecard

    Comment and details

    The strongest GLM, at about US$ 2 a run. Its first run fixed all four flaws the probes cover, the CSV formula injection among them; the third changed two business values and fell to 541.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 635 9 of 13 flaws fixed 1 false positive contract kept 22/22 checks 15min US$ 1.69 Scorecard
    • Run 2 · 628 8 of 13 flaws fixed no false positive contract kept 22/22 checks 26min US$ 2.13 Scorecard
    • Run 3 · 541 7 of 13 flaws fixed no false positive 2 business values changed 22/22 checks 17min US$ 1.77 Scorecard
  15. 15
    GPT-5.6-terra OpenAI · effort xhigh 3 of 3 runs (625 · 611 · 645)
    625 of 1000 LEB Silver
    • Security 211/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 35/50

    Penalties: none Discovery 45.8 Brier 0.000 Scorecard

    Comment and details

    Reports fewer flaws than it fixes: its first run listed 7 and fixed 9, the rest done in passing without a word in the report. Compatibility held in two of three runs, which keeps it 15th despite the short report.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 625 9 of 13 flaws fixed no false positive contract kept 22/22 checks 8min Scorecard
    • Run 2 · 611 10 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 8min Scorecard
    • Run 3 · 645 10 of 13 flaws fixed no false positive contract kept 22/22 checks 7min Scorecard
  16. 16
    GLM-5.3-Flash Z.AI · effort high 1 of 3 runs · not official no published cutoff or earlier release
    624 of 1000 LEB Silver
    • Security 211/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 34/50

    Penalties: none Discovery 70.8 Brier 0.016 Cost US$ 0.08 a run Scorecard

    Comment and details

    One run so far: 624 for US$ 0.08, about a tenth of GLM-5.3's cost for a slightly higher score. It left the CSV formula injection and the hardcoded secrets in place.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 624 8 of 13 flaws fixed no false positive contract kept 22/22 checks 15min US$ 0.08 Scorecard
  17. 17
    GLM-5.3 Z.AI · effort high 3 of 3 runs (629 · 621 · 604) no published cutoff or earlier release
    621 of 1000 LEB Silver
    • Security 203/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 39/50

    Penalties: none Discovery 70.8 Brier 0.014 Cost US$ 0.57 a run Scorecard

    Comment and details

    Steady (604 to 629) and cheap (US$ 0.47 to 0.73 a run), with compatibility at 100 in all three runs. It never fixed MD5 or the CSV formula injection.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 629 8 of 13 flaws fixed no false positive contract kept 22/22 checks 17min US$ 0.73 Scorecard
    • Run 2 · 621 7 of 13 flaws fixed no false positive contract kept 22/22 checks 15min US$ 0.51 Scorecard
    • Run 3 · 604 7 of 13 flaws fixed no false positive contract kept 22/22 checks 13min US$ 0.47 Scorecard
  18. 18
    DeepSeek V4 Flash DeepSeek · effort xhigh 1 of 3 runs · not official no published cutoff or earlier release
    617 of 1000 LEB Silver
    • Security 207/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 31/50

    Penalties: none Discovery 58.3 Brier 0.004 Cost US$ 0.18 a run Scorecard

    Comment and details

    One run so far, through Novita: 617 for US$ 0.18, against 282 for the same model at high. It fixed 7 flaws and left the CSV injection, MD5 and the hardcoded secrets in place.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 617 7 of 13 flaws fixed no false positive contract kept 22/22 checks 28min US$ 0.18 Scorecard
  19. 19
    GPT-6.1-sol OpenAI · effort ultra 3 of 3 runs (597 · 656 · 616)
    616 of 1000 LEB Silver
    • Security 228/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 39/50

    Penalties: none Discovery 70.8 Brier 0.001 Scorecard

    Comment and details

    Ultra effort scored 45 points below GPT-6.1-sol at xhigh, across runs from 597 to 656. With three subagents its first run kept MD5; with two, its second fixed 10 flaws, the CSV formula injection and MD5 among them.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance
    Zero
    Architecture, Clean code
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 597 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 24min Scorecard
    • Run 2 · 656 10 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 13min Scorecard
    • Run 3 · 616 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 11min Scorecard
  20. 20
    GPT-5.6-sol OpenAI · effort xhigh 3 of 3 runs (612 · 612 · 608)
    612 of 1000 LEB Silver
    • Security 246/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 32/50

    Penalties: -15 Discovery 70.8 Brier 0.001 Scorecard

    Comment and details

    Security at 246 of 250 and the CSV formula injection fixed, but the fix also turned the '-' of a ticket with no technician into "'-", a new bug, and every run changed one business value (−30). The most repeatable GPT: 608 to 612.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance
    Zero
    Architecture, Clean code
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 612 10 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 7min Scorecard
    • Run 2 · 612 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 15min Scorecard
    • Run 3 · 608 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 15min Scorecard
  21. 21
    DeepSeek V4.1 Flash DeepSeek · effort high 3 of 3 runs (625 · 612 · 597) no published cutoff or earlier release
    612 of 1000 LEB Silver
    • Security 203/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 30/50

    Penalties: none Discovery 58.3 Brier 0.013 Cost US$ 0.02 a run Scorecard

    Comment and details

    The cheapest agent here, US$ 0.01 to 0.04 a run, two of them under two minutes, for 597 to 625. Compatibility held in all three; the lowest explanation score among the agents above 600 (30 of 50).

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 625 9 of 13 flaws fixed no false positive contract kept 22/22 checks 9min US$ 0.04 Scorecard
    • Run 2 · 612 8 of 13 flaws fixed no false positive contract kept 22/22 checks 1min US$ 0.01 Scorecard
    • Run 3 · 597 7 of 13 flaws fixed no false positive contract kept 22/22 checks 1min US$ 0.02 Scorecard
  22. 22
    GPT-5.6-luna OpenAI · effort xhigh 3 of 3 runs (599 · 601 · 624)
    601 of 1000 LEB Silver
    • Security 220/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 32/50

    Penalties: none Discovery 58.3 Brier 0.000 Scorecard

    Comment and details

    The smallest GPT-5.6, within 24 points of terra and sol across 599 to 624. It changed one business value in each run (−30).

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance
    Zero
    Architecture, Clean code
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 599 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 12min Scorecard
    • Run 2 · 601 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 10min Scorecard
    • Run 3 · 624 10 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 14min Scorecard
  23. 23
    GLM-5.3-FlashX Z.AI · effort high 1 of 3 runs · not official no published cutoff or earlier release
    597 of 1000 LEB Bronze
    • Security 185/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 33/50

    Penalties: none Discovery 70.8 Brier 0.015 Cost US$ 0.10 a run Scorecard

    Comment and details

    One run so far: 597 for US$ 0.10, 27 points below GLM-5.3-Flash. It fixed 7 flaws and left the CSV injection, MD5 and the hardcoded secrets in place.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 597 7 of 13 flaws fixed no false positive contract kept 22/22 checks 9min US$ 0.10 Scorecard
  24. 24
    Gemini 3.8 Flash Google · effort high 3 of 3 runs (687 · 588 · 100) no published cutoff or earlier release
    588 of 1000 LEB Bronze
    • Security 181/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 28/50

    Penalties: none Discovery 58.3 Brier 0.003 Cost US$ 0.71 a run Scorecard

    Comment and details

    Three very different runs: 687 (it would have placed 6th), 588, and 100 from a run that stopped after ten minutes on a database server it started itself. The two complete runs kept MD5 and both secrets and left the CSV export serving every client.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority
    Run by run
    • Run 1 · 687 8 of 13 flaws fixed no false positive contract kept 22/22 checks 13min US$ 1.00 Scorecard
    • Run 2 · 588 7 of 13 flaws fixed no false positive contract kept 22/22 checks 25min US$ 0.85 Scorecard
    • Run 3 · 100 0 of 13 flaws fixed no false positive contract kept 22/22 checks 10min US$ 0.28 Scorecard
  25. 25
    Gemini 3.7 Flash Google · effort high 3 of 3 runs (587 · 587 · 556) no published cutoff or earlier release
    587 of 1000 LEB Bronze
    • Security 181/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 27/50

    Penalties: none Discovery 58.3 Brier 0.002 Cost US$ 0.60 a run Scorecard

    Comment and details

    The generation before Gemini 3.8 Flash, one point behind it and far steadier (556 to 587). Its first run is the only one outside Claude to flatten the nested ifs; its second searched the VM for the answer key, which is not there.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 587 7 of 13 flaws fixed no false positive contract kept 22/22 checks 12min US$ 0.72 Scorecard
    • Run 2 · 587 7 of 13 flaws fixed no false positive contract kept 22/22 checks 16min US$ 0.60 Scorecard
    • Run 3 · 556 7 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 7min US$ 0.49 Scorecard
  26. 26
    GPT-5.5 OpenAI · effort xhigh 3 of 3 runs (558 · 568 · 536)
    558 of 1000 LEB Bronze
    • Security 177/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 32/50

    Penalties: none Discovery 58.3 Brier 0.006 Scorecard

    Comment and details

    The previous GPT generation, 43 to 108 points below the GPT-5.6 and GPT-6 models: 8 flaws found and 7 fixed in every run, MD5 and the CSV formula injection always left in place.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 2 · 558 7 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 6min Scorecard
    • Run 3 · 568 7 of 13 flaws fixed no false positive contract kept 22/22 checks 5min Scorecard
    • Run 4 · 536 7 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 5min Scorecard
  27. 27
    Gemini 3.8 Flash Google · effort medium 1 of 3 runs · not official no published cutoff or earlier release
    550 of 1000 LEB Bronze
    • Security 147/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 24/50

    Penalties: none Discovery 45.8 Brier 0.201 Cost US$ 0.19 a run Scorecard

    Comment and details

    One run so far, in 4min 34s for US$ 0.19: 550. It fixed 6 flaws and claimed SQL injection in two functions that only take integers.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; Secrets hardcoded in the config; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 550 6 of 13 flaws fixed 2 false positives contract kept 22/22 checks 5min US$ 0.19 Scorecard
  28. 28
    Qwen3 Coder Next Alibaba (Qwen team) · default effort (not configurable) 1 of 3 runs · not official
    507 of 1000 LEB Bronze
    • Security 125/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 18/50

    Penalties: -15 Discovery 54.2 Brier 0.301 Cost US$ 1.53 a run Scorecard

    Comment and details

    The weakest explanation (18 of 50) and the worst calibration here: three flaws that cannot exist reported at confidence 100, and no tests run. It is, on the other hand, the only one of the bottom nine that fixed the N+1 query.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Performance, Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Reflected XSS in the search; Formula injection in the CSV export; Unsalted MD5 passwords; Secrets hardcoded in the config; Any ticket readable by id (IDOR); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 507 5 of 13 flaws fixed 3 false positives contract kept 22/22 checks 16min US$ 1.53 Scorecard
  29. 29
    DeepSeek V4 Pro DeepSeek · effort high 3 of 3 runs (604 · 496 · 432) no published cutoff or earlier release
    496 of 1000 LEB Bronze
    • Security 181/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 56/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 30/50

    Penalties: none Discovery 58.3 Brier 0.026 Cost US$ 0.16 a run Scorecard

    Comment and details

    The larger DeepSeek: all three runs (604, 496 and 432) score below DeepSeek V4.1 Flash's official 612. The first rated SQL injection at confidence 100 in two functions that only take integers; the other two left the N+1 query in place.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Compatibility
    Zero
    Architecture, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Unsalted MD5 passwords; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 604 8 of 13 flaws fixed 2 false positives contract kept 22/22 checks 8min US$ 0.03 Scorecard
    • Run 2 · 496 6 of 13 flaws fixed no false positive contract kept 22/22 checks 23min US$ 0.42 Scorecard
    • Run 3 · 432 6 of 13 flaws fixed 1 false positive contract kept 22/22 checks 12min US$ 0.04 Scorecard
  30. 30
    MiniMax-M3 MiniMax · effort thinking 3 of 3 runs (462 · 616 · 434)
    462 of 1000 LEB Bronze
    • Security 220/250
    • Architecture 0/200
    • Bugs 150/150
    • Performance 0/150
    • Clean code 25/100
    • Compatibility 40/100
    • Explanation 27/50

    Penalties: none Discovery 62.5 Brier 0.021 Cost US$ 0.12 a run Scorecard

    Comment and details

    Fixed 8 flaws in every run, with bugs at full marks, but changed business values (compatibility at 40 in its official run) and left the N+1 query. Its second run scored 616; its third broke four of the 22 characterization checks.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Bugs
    Zero
    Architecture, Performance
    Never fixed, in any run
    Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority
    Run by run
    • Run 1 · 462 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 4min US$ 0.12 Scorecard
    • Run 2 · 616 8 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 6min US$ 0.10 Scorecard
    • Run 3 · 434 8 of 13 flaws fixed 1 false positive 1 business value changed 18/22 checks 5min US$ 0.13 Scorecard
  31. 31
    Kimi K2.7 Code Moonshot AI · default effort (not configurable) 3 of 3 runs (415 · 415 · 512)
    415 of 1000 LEB Bronze
    • Security 203/250
    • Architecture 0/200
    • Bugs 86/150
    • Performance 0/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 26/50

    Penalties: none Discovery 50.0 Brier 0.225 Cost US$ 0.51 a run Scorecard

    Comment and details

    Left the N+1 query in every run and reported flaws that do not exist in all three, SQL injection in functions that only take integers among them. Its third run, 512, fixed one flaw more than the other two.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Compatibility
    Zero
    Architecture, Performance, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Unsalted MD5 passwords; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 415 6 of 13 flaws fixed 2 false positives contract kept 22/22 checks 14min US$ 0.48 Scorecard
    • Run 2 · 415 6 of 13 flaws fixed 1 false positive contract kept 22/22 checks 4min US$ 0.25 Scorecard
    • Run 3 · 512 7 of 13 flaws fixed 2 false positives contract kept 22/22 checks 15min US$ 0.81 Scorecard
  32. 32
    GPT-5.3-Codex OpenAI · effort xhigh 1 of 3 runs · not official
    403 of 1000 LEB Bronze
    • Security 147/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 0/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 27/50

    Penalties: none Discovery 37.5 Brier 0.015 Cost US$ 0.54 a run Scorecard

    Comment and details

    One run so far, through opencode: 403, the lowest GPT. It found 6 flaws and fixed 5, leaving the N+1 query, the session fixation at login and MD5 in place, with no false positive.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Compatibility
    Zero
    Architecture, Performance, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; Secrets hardcoded in the config; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 403 5 of 13 flaws fixed no false positive contract kept 22/22 checks 7min US$ 0.54 Scorecard
  33. 33
    Kimi K2.7 Code (highspeed) Moonshot AI · default effort (not configurable) 3 of 3 runs (402 · 363 · 402) no published cutoff or earlier release
    402 of 1000 LEB Bronze
    • Security 155/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 0/150
    • Clean code 25/100
    • Compatibility 70/100
    • Explanation 23/50

    Penalties: none Discovery 54.2 Brier 0.216 Cost US$ 0.74 a run Scorecard

    Comment and details

    The fast variant of Kimi K2.7 Code, 13 points below it: 5 flaws fixed in every run, two flaws that do not exist reported in each, and the N+1 query always left in place.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Zero
    Architecture, Performance
    Never fixed, in any run
    Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; Secrets hardcoded in the config; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 402 5 of 13 flaws fixed 2 false positives 1 business value changed 22/22 checks 2min US$ 0.38 Scorecard
    • Run 2 · 363 5 of 13 flaws fixed 2 false positives 1 business value changed 22/22 checks 7min US$ 0.68 Scorecard
    • Run 3 · 402 5 of 13 flaws fixed 2 false positives contract kept 22/22 checks 10min US$ 1.16 Scorecard
  34. 34
    GLM-5.2 Z.AI · effort high 1 of 3 runs · not official
    388 of 1000 Failed
    • Security 172/250
    • Architecture 0/200
    • Bugs 86/150
    • Performance 0/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 30/50

    Penalties: none Discovery 50.0 Brier 0.001 Cost US$ 0.35 a run Scorecard

    Comment and details

    One run so far: 388, below the pass line and 233 points under GLM-5.3. It left the N+1 query, the leaked file handle and the session fixation in place; what it did report was accurate, with no false positive.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Compatibility
    Zero
    Architecture, Performance, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; File handle leaked on the error path; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 388 5 of 13 flaws fixed no false positive contract kept 22/22 checks 6min US$ 0.35 Scorecard
  35. 35
    Nex N2.5 Pro Nex AGI · effort high 2 of 3 runs (317 · 428) · not official no published cutoff or earlier release
    317 of 1000 Failed
    • Security 194/250
    • Architecture 0/200
    • Bugs 107/150
    • Performance 75/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 31/50

    Penalties: -160 Discovery 62.5 Brier 0.000 Cost US$ 3.84 a run Scorecard

    Comment and details

    Too slow to finish three runs: 10h 03min and 19h 35min, and a third voided after more than 21 hours when the operator's connection failed, so it stays at two and publishes the lower. Both runs were strong on security and both broke the contract: the first made every function return nothing without a logged-in user (8 checks, −160, 317), the second made the CSV export throw for a caller that has already printed output (3 checks, 428).

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Zero
    Architecture, Clean code
    Never fixed, in any run
    A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 317 9 of 13 flaws fixed no false positive 1 business value changed 14/22 checks 10h 3min US$ 2.80 Scorecard
    • Run 2 · 428 10 of 13 flaws fixed no false positive 2 business values changed 19/22 checks 19h 35min US$ 4.87 Scorecard
  36. 36
    Claude Haiku 4.5 Anthropic · default effort (not configurable) 3 of 3 runs (317 · 369 · 232)
    317 of 1000 Failed
    • Security 138/250
    • Architecture 0/200
    • Bugs 86/150
    • Performance 0/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 23/50

    Penalties: none Discovery 37.5 Brier 0.215 Cost US$ 0.29 a run Scorecard

    Comment and details

    The small Claude, at US$ 0.21 to 0.40 a run, below the pass line in all three. Its first run's visibility fix hides a client's own tickets, comparing an integer with the string mysqli returns; its third broke three characterization checks.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Zero
    Architecture, Performance, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; Session fixation at login; Unsalted MD5 passwords; File handle leaked on the error path; One query per ticket for the technician (N+1); A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 317 4 of 13 flaws fixed 2 false positives 1 business value changed 22/22 checks 3min US$ 0.21 Scorecard
    • Run 2 · 369 5 of 13 flaws fixed 2 false positives contract kept 22/22 checks 7min US$ 0.40 Scorecard
    • Run 3 · 232 5 of 13 flaws fixed 1 false positive 1 business value changed 19/22 checks 10min US$ 0.27 Scorecard
  37. 37
    DeepSeek V4 Flash DeepSeek · effort high 2 of 3 runs (612 · 282) · not official no published cutoff or earlier release
    282 of 1000 Failed
    • Security 99/250
    • Architecture 0/200
    • Bugs 96/150
    • Performance 0/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 22/50

    Penalties: -35 Discovery 50.0 Brier 0.122 Cost US$ 0.04 a run Scorecard

    Comment and details

    Two runs that could hardly differ more: the first fixed 9 flaws for 612, the second's SQL-injection fix throws on every search and its client CSV comes out on one line, 282. It publishes the lower; which weights the hosts served under this name is not certain.

    A written reading drawn from the scorecards and verdicts; it is not part of the score.

    Full marks
    Compatibility
    Zero
    Architecture, Performance, Clean code
    Never fixed, in any run
    Formula injection in the CSV export; A dispatcher that does everything; Magic numbers for status and priority; Four levels of nested ifs
    Run by run
    • Run 1 · 612 9 of 13 flaws fixed no false positive 1 business value changed 22/22 checks 9min US$ 0.07 Scorecard
    • Run 2 · 282 3 of 13 flaws fixed 1 false positive contract kept 21/22 checks 7min US$ 0.00 Scorecard

Flaw by flaw

What each agent found and fixed among the planted flaws. The hard ones are flaws of absence — a missing authorization check, a session never regenerated, a file left open on the error path.

The table shows the top 10 of 37 agents. Every agent's result, flaw by flaw, is in its scorecard.

  • fixed
  • found, not fixed
  • missed
Flaw Claude Sonnet 5.5 · xhigh Claude Sonnet 5.5 · max Claude Sonnet 5.5 · max (ultracode) Claude Fable 5.1 Claude Opus 5.5 GPT-6-astra · ultra GPT-6.1-sol · xhigh GPT-6.1-sol pro Grok 4.7 · high Grok 4.6
SQL injection in the search SEC-001 · critical · easy 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed
Reflected XSS in the search SEC-003 · high · easy 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed
Formula injection in the CSV export SEC-008 · medium · hard 6/6 fixed 6/6 fixed 6/6 fixed 2/6 found, not fixed 2/6 found, not fixed 6/6 fixed 6/6 fixed 2/6 found, not fixed 6/6 fixed 0/6 missed
Session fixation at login SEC-013 · high · hard 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed
Unsalted MD5 passwords SEC-014 · high · easy 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 3/8 found, not fixed 8/8 fixed 8/8 fixed 3/8 found, not fixed 3/8 found, not fixed
Secrets hardcoded in the config SEC-015 · high · easy 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 3/8 found, not fixed 6/8 found, not fixed
Any ticket readable by id (IDOR) SEC-017 · critical · hard 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 9/10 fixed 9/10 fixed 9/10 fixed 10/10 fixed 10/10 fixed
Division by zero in the SLA average BUG-001 · high · moderate 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed
File handle leaked on the error path BUG-004 · medium · hard 5/6 fixed 6/6 fixed 4/6 fixed 5/6 fixed 5/6 fixed 4/6 fixed 4/6 fixed 5/6 fixed 5/6 fixed 4/6 fixed
One query per ticket for the technician (N+1) PERF-001 · high · moderate 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed
A dispatcher that does everything ARCH-002 · high · easy 2/10 found, not fixed 1/10 found, not fixed 0/10 missed 0/10 missed 4/10 found, not fixed 2/10 found, not fixed 2/10 found, not fixed 0/10 missed 0/10 missed 0/10 missed
Magic numbers for status and priority ARCH-009 · low · moderate 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed
Four levels of nested ifs CLN-007 · medium · easy 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 0/8 missed 2/8 found, not fixed 0/8 missed 2/8 found, not fixed 0/8 missed 2/8 found, not fixed

What stood out

  • Six agents fixed the formula injection in the CSV (SEC-008) — Sonnet 5.5 in all three of its settings, GPT-6.1-sol, GPT-5.6-sol and Grok 4.7 — with the fix the answer key expects; Sonnet 5.5 is the only model with security at 250 of 250, in all three of its settings. GPT-5.6-sol also turned the - of a ticket with no technician into '-, a new bug (−15).
  • Architecture was the weakest category for everyone: 50 of 200 for Opus 5.5, 25 for Sonnet 5.5 at xhigh and for GPT-6.1-sol, 12 for Sonnet 5.5 at max, and 0 for the other twenty-eight. Nobody split the dispatcher that does everything; those three named it and declined to restructure it.
  • Two published runs broke the contract mechanically. All thirty-seven stayed on mysqli and reported no decoy, and thirty-five kept the 22 characterization checks green; DeepSeek V4 Flash publishes a run whose search throws, and Nex N2.5 Pro one whose functions return nothing without a logged-in user (8 checks broken, −160). Judgement is what separated the rest: twenty-four kept compatibility at 100, while the other thirteen each changed at least one business value (−30 each), most often the SLA average, scoped to each client.
  • Twenty-six scores are official: Claude Sonnet 5.5 at xhigh, 809 (runs of 825, 809 and 724), Claude Sonnet 5.5 at max, 807 (807, 773 and 820), Claude Sonnet 5.5 at max in multi-agent mode, 774 (774, 820 and 759), Claude Opus 5.5, 717 (711, 717 and 805), GPT-6.1-sol, 661 (666, 653 and 661), GPT-6.1-sol at ultra, 616 (597, 656 and 616), Grok 4.7, 638 (638, 663 and 607), Grok 4.6, 633 (633, 640 and 620), Grok 4.7 at xhigh, 631 (640, 617 and 631), GPT-6-astra at ultra, 666 (668, 666 and 651), Kimi K3 at high, 629 (629, 634 and 574), GPT-6-astra, 628 (661, 596 and 628), GPT-5.6-terra, 625 (625, 611 and 645), GLM-5.3 Prime, 628 (635, 628 and 541), GLM-5.3, 621 (629, 621 and 604), DeepSeek V4.1 Flash, 612 (625, 612 and 597), GPT-5.6-sol, 612 (612, 612 and 608), GPT-5.6-luna, 601 (599, 601 and 624), Gemini 3.8 Flash at high, 588 (687, 588 and 100), GPT-5.5, 558 (558, 568 and 536), DeepSeek V4 Pro, 496 (604, 496 and 432), MiniMax-M3 with the thinking variant, 462 (462, 616 and 434), Kimi K2.7 Code, 415 (415, 415 and 512), Kimi K2.7 Code highspeed, 402 (402, 363 and 402), Gemini 3.7 Flash at high, 587 (587, 587 and 556), and Claude Haiku 4.5, 317 (317, 369 and 232), each the median of three runs. A single run can sit more than 100 points from the median: Opus's third scored 805, Sonnet's third 724, DeepSeek V4 Pro's first 604, Gemini 3.8 Flash's third, which stopped after ten minutes, 100, and astra's first, 661, had placed it 5th, from judgement calls such as which flaws to leave unfixed. Claude Fable 5.1 has two runs, 781 and 764, and publishes the lower, 4th
  • Gemini 3.8 Flash at high is official at 588, 24th, Bronze, the median of three very different runs (687, 588 and 100). Gemini 3.7 Flash, the generation before, is official at 587 (587, 587 and 556), 25th. Its first run, 6th, fixed 8 of the 13 planted flaws and is the only one outside Claude to flatten the nested ifs (CLN-007); its second fixed 7, and searched the VM for the answer key, by name and by the hash quoted in the task. The key is not on the VM, and the search turned up nothing the run did not already have. Its third stopped after ten minutes, on a database server it started that kept its own shell command from ever returning: no report, the code untouched, scored as delivered. The two complete runs kept MD5 and both secrets and left the CSV export serving every client. At medium effort the same model scored 550 (27th) in 4min 34s: it fixed 6 flaws and claimed SQL injection in two functions that only take integers.
  • GPT-6-astra at ultra is the strongest GPT agent here, official at 666 (runs of 668, 666 and 651), 6th, with two or three subagents in 10 to 15 minutes. GPT-6.1-sol follows, official at 661 (666, 653 and 661), 7th, and GPT-6.1-sol pro at 654, 8th, from one run. GPT-6-astra at xhigh, official at 628, is 13th. Its first run fixed the CSV injection and kept MD5; its official run did the opposite, migrating MD5 to password_hash and leaving the CSV injection alone; GPT-6.1-sol's official run fixed both
  • Multi-agent work matched one agent at best, at many times the cost. Claude Sonnet 5.5 in Claude Code's multi-agent mode, at max effort, scored 774 in its first run (7 workflows, 68 subagents, 7h 55min, US$ 233), 820 in its second (2 workflows, 105 subagents, 2h, US$ 171) and 759 in its third (3 workflows, 145 subagents, 3h 34min, US$ 119). It is official at 774, 3rd, Gold. The same model at the same max effort as a single agent took about half an hour a run and is official at 807 (runs of 807, 773 and 820), 2nd; at xhigh it is official at 809. Its third report scored 47 of 50, the best explanation here, and no run touched the architecture. In each run a few subagents (nine, eight, then seven) fell back to an older Sonnet after a safety classifier stopped them; none of any delivery came from them.
  • GPT-6.1-sol at ultra is official 45 points below itself at xhigh (616 against 661), the median of three runs spread over 59 points (597, 656 and 616). With three subagents, its first found 9 of the 13 planted flaws, kept MD5 and did not name the dispatcher; with two, its second fixed 10, the CSV injection and MD5 among them. Each took under half an hour, not the 8 hours of the multi-agent Sonnet.
  • Grok 4.7 is the strongest model here from outside Anthropic and OpenAI (638, 9th, Silver, official over three runs of 638, 663 and 607), with an explanation at the level of the best GPT reports (42 of 50). In its first run it is also the only agent that used the web, to read the PHP manual page for fputcsv; the VM blocks GitHub, not the web. At xhigh, Grok 4.7 is official at 631 (640, 617 and 631), 11th, below its own high effort, with reports scored 42 to 44 of 50, the best explanations from outside Anthropic and OpenAI. Grok 4.6 follows, official at 633 (10th; runs of 633, 640 and 620), for about an eighth of the cost (US$ 0.31 against 2.43).
  • GLM-5.3 Prime is the strongest GLM model (628, 14th, Silver), official over three runs through two hosts (635, 628 and 541): the first fixed all four probe-covered flaws, the CSV injection among them, for US$ 1.69; the second left the injection alone. GLM-5.3 is official at 621 (17th), the median of three runs (629, 621 and 604), the last two served by another host; GLM-5.3-Flash scores 624 for about a tenth of the cost, and GLM-5.3-FlashX 597
  • DeepSeek V4.1 Flash is the cheapest agent here (US$ 0.01 to 0.04 a run; its last two took under two minutes each) and is official at 612 (runs of 625, 612 and 597), 21st. DeepSeek V4 Flash at xhigh, through Novita, scored 617 in its first run (18th). At high, run through two other hosts, it publishes 282, the lower of its two runs (612 and 282), last: the first fixed eight flaws fully, while the second's SQL-injection fix throws on every search and its client CSV comes out on one line. DeepSeek now routes the V4 Flash name to V4.1 on its own API, so which weights those hosts served is not certain. DeepSeek V4 Pro, the larger model, is official at 496, the median of three runs through two hosts (604, 496 and 432), all below DeepSeek V4.1 Flash's official 612: the first rates SQL injection at confidence 100 in two functions that only take integers, and the other two left the N+1 query in place.
  • Twenty-four agents did not run at xhigh. Kimi K2.7 Code (in both of Moonshot's ids), Qwen3 Coder Next and Claude Haiku 4.5 have no effort setting and ran at their model's default; the two Grok models, the five GLM models, the three DeepSeek models and Gemini 3.8 and 3.7 Flash and Nex N2.5 Pro ran at high, in opencode, and Gemini 3.8 Flash also at medium; MiniMax-M3 ran in opencode with its thinking variant and Kimi K3 with its high variant; Sonnet 5.5 ran twice at max, once in multi-agent mode, and GPT-6.1-sol and GPT-6-astra once each at ultra. The rest ran at xhigh.
  • Qwen3 Coder Next, DeepSeek V4 Pro, MiniMax-M3 with the thinking variant, Kimi K2.7 Code, GPT-5.3-Codex, Kimi K2.7 Code highspeed, GLM-5.2, Nex N2.5 Pro, Claude Haiku 4.5 and DeepSeek V4 Flash close the table (507, 496, 462, 415, 403, 402, 388, 317, 317 and 282; the last four are below the pass line). Nex N2.5 Pro is too slow to finish three runs: 10h 03min and 19h 35min, the longest single-agent runs here, and a third voided after more than 21 hours, so it stays at two; both fixed most of the security flaws, both broke the contract, and it publishes the lower, 317 (317 and 428). Claude Haiku 4.5 is official at 317 (runs of 317, 369 and 232); in the first, 2.5 minutes and US$ 0.21, its visibility fix hides a client's own tickets on the main page, comparing an integer with the string mysqli returns. All but Qwen3 Coder Next left the N+1 query in place. Qwen3 Coder Next has the weakest explanation (18 of 50) and the worst calibration (Brier 0.301): it reported three flaws that cannot exist at confidence 100, SQL injection in two functions that only take integers among them, as Kimi K2.7 Code did. It also ran no tests.
  • Places 9 to 23 sit within 41 points (638 to 597), inside the noise of a single run. GPT-5.6-terra reported the fewest planted flaws and still ranks 15th, on compatibility and on what it did fix.
  • The GPT models are among the best calibrated (Brier 0.015 or less): fewer findings, each stated with high confidence and nearly all real.

Finding is not fixing

Side by side, the three Claude models find nearly the same flaws and fix different numbers of them. One total hides which of the two it is measuring.

  1. They find nearly the same. In the run that counts for each, Sonnet 5.5 at xhigh reported 12 of the 13 planted flaws and Fable 5.1 and Opus 5.5 11; across their eight runs, each found 11 or 12
  2. Sonnet 5.5 fixes more, in the run that counts. Its official run fixed 11 of the 13, against 9 for Opus's official run and 10 for Fable's. Run by run it varies: Sonnet fixed 11, 11 and 8, Opus 9 in each of its three, Fable 9 and 10. Fable and Opus found and explained the formula injection in the CSV and left it in place, to protect the file's consumers; Sonnet fixed it in two of its three runs, prefixing only the cells that start a formula. Its 92-point lead over Opus's official 717 comes from clean code (+100), not from security (+17); its 45-point lead over Fable's 764 is security (+17) and architecture (+25)
  3. More compute did not change the pattern. The single-agent Claude runs took 16 to 27 minutes each, and the three at the same max effort scored 807, 773 and 820. Sonnet 5.5 in multi-agent mode took 7h 55min, 2h and 3h 34min and US$ 233, 171 and 119 for 774, 820 and 759, and is official at 774 against the single agent's 807: at best it tied the single-agent 820, with the same fixes and a deeper check of its own code, and in none of its three runs did the dispatcher reach its report.
  4. How to read it. On this task the three Claude models see the same problems; what separates them is how far they go in fixing them within the contract. In the median, Sonnet went further; a single run can reverse it, as Sonnet's third did. One task, two or three runs per model: a pattern worth testing, not a verdict

Read this before quoting a number

  • Up to three runs per agent. An official LEB score is the median of three independent runs. Until an agent has three, its score is the lower of its totals so far, the totals are listed next to it, and everything shown with the score comes from that same run.
  • The judge is an AI. Claude Opus 5.5 applied the published rubric to every delivery without knowing which model wrote it — each was anonymised — and the explanation was scored by a separate judge that saw neither the answer key nor the other scores.
  • The judge is also a contestant. Claude Opus 5.5 is one of the agents evaluated, and six of the thirty-seven are Claude models, Sonnet 5.5 three times: they hold the top five places and the second to last. Anonymity limits that bias; it does not remove it, since a model can recognise its own style. Every verdict is published with its rationale, flaw by flaw, and the four verdicts changed in review say why; one of them adds 41 points to GPT-5.6-terra's total.
  • Twenty-one attempts were left unscored, each before any judge saw it. Twice, Claude Fable 5.1's safeguards stopped one of its responses while it worked on the security flaws, and its client handed the rest of the run to Claude Opus 4.8, which wrote the whole report; a run must come from one model, so both are void, and Fable 5.1 shows its two complete runs. Four ended on the provider's side before the delivery was complete: GPT-5.6-sol pro when its gateway ran out of credit, and Gemini 3.8 Flash three times, on a gateway timeout and twice on a rate limit. Five were set up wrong: the first multi-agent Sonnet 5.5 run was stopped to give the machine more processors, the second found its client logged out and never reached the model, as did Claude Fable 5.1's third, Claude Haiku 4.5's first attempt had its client mode changed mid-run, and a Gemini attempt was started outside the task's folder. One finished, but GPT-5.6-terra ran in a different client from its first run, and was repeated in the same one. Six finished after their agent already had its three runs, two of GPT-6.1-sol at xhigh, three at ultra and one of Gemini 3.7 Flash; a fourth run would let the published score be chosen. Two finished, a Grok 4.7 run at xhigh and a Kimi K3 run at high, but their VMs were restored before the deliveries were copied off them, so nothing was left to score. One was cut off by the operator's own connection: Nex N2.5 Pro's third run, after more than 21 hours, which the operator chose not to repeat because the model is too slow. All twenty-one are kept in the repository.
  • The answer key is public. The failure matrix of LEB-100-A has been in the public repository since 13 July 2026. The runs of 29 September had the VM's name block only, so no agent could fetch it by accident; the runs of 30 September also had GitHub's addresses blocked, and the session logs of every run that left one show no request to GitHub at all. What a model saw in training depends on how far its training data reaches: each run records the cutoff its provider publishes. Without one, the provider's release date bounds it, since a model cannot train on data from after its release. A model whose cutoff or release is later, or that has neither, is marked in the leaderboard.
  • The benchmark's own tests were fixed. Scoring these runs exposed two defects in the evaluation tooling: an SQL loader that split a statement on a semicolon inside a comment, and CSV checks that read a temporary file the contract never promised. Both were fixed before scoring, the same way for every agent, and are recorded in the repository.
  • The machine changed on 30 September. Until then the agents ran as the VM's administrator, with sudo, on a machine that kept earlier runs' leftovers. Their session logs show no agent reading another's work; the opencode agents shared a leftover test database that held only the seed rows. From Qwen3 Coder Next on, runs are made as an unprivileged user on a machine cleaned before each run.
  • Not every run parameter was recorded. The exact model version and the temperature were not, and cost and model time only where the client kept them; the three runs that left no session log at all (GPT-5.5's first, MiniMax-M3's at its default and Kimi K3's) were withdrawn on 4 October 2026 and are not published. MiniMax-M3 at its default is gone from the table; Kimi K3 is back with a recorded run at high. Each run marks what is missing instead of guessing.

Audit it

Every delivery, mechanical report, verdict and scorecard is in the repository, next to the specification that produced them. The same data is here as two spreadsheets: one row per run, and one per run and planted flaw.