Skip to content
AI Benchmark · LEB

Can an AI maintain legacy code without breaking it?

LEB — the LLM Engineering Benchmark — hands an AI agent a legacy system in production, with flaws planted in it and consumers that depend on how it behaves today. It measures the work that dominates real engineering: finding the flaws, fixing them, keeping every contract intact, and explaining the decisions like a senior engineer would.

Why another benchmark

Most benchmarks measure code written from scratch, or one isolated issue solved. Neither is what most engineering is: evolving a system that other people already depend on. LEB scores security, architecture, bugs, performance, clean code, compatibility and the quality of the explanation — and it takes points away from the agent that rewrites everything, swaps technologies without need, or breaks a public contract.

Rewriting from scratch is not engineering. It is running away.

How a run works

  1. 01

    A legacy system, with planted flaws

    The agent receives the code, a manifest of its public surface — the contract — and a neutral task: report the problems, fix what should be fixed, keep compatibility, justify every decision. It is never told which flaws exist, how many, or where.

  2. 02

    The agent works alone

    In mode A it gets tools and a budget of turns; in mode S, one prompt and one answer. It hands back the changed code, a technical report, and an index of its findings, each with a 0–100 confidence.

  3. 03

    Machines check the code

    Characterization tests run on the legacy code and on the delivery: public behaviour that changed is a regression. Probes then attack each fixable flaw — the injection payload, the empty dataset, the query counter — and report whether it is still there.

  4. 04

    A judge checks the report

    Each finding is matched against the Official Failure Matrix, a hidden answer key that also holds decoys: plausible flaws that do not exist, and cost points when reported. A second judge scores the explanation blind. A deterministic scorer turns it all into 0–1000.

1000 points, and how they are lost

Every instance is worth exactly 1000: the raw points of each category are normalised to its weight, so scores compare across instances of a level.

Compatibility starts at 100 and only goes down. Migrating mysqli to PDO without need costs 20; changing a public signature costs 30, per function.

Global penalties come off the total: a new bug −15, each broken characterization test −20, a needless rewrite −25, each decoy reported −5.

  • Security SEC 250
  • Architecture ARCH 200
  • Bugs BUG 150
  • Performance PERF 150
  • Clean code CLN 100
  • Compatibility COMP 100
  • Explanation EXPL 50

Grades

  • LEB Platinum 900–1000 · ready for critical legacy
  • LEB Gold 750–899 · solid engineering
  • LEB Silver 600–749 · useful with supervision
  • LEB Bronze 400–599 · needs a full review
  • Failed < 400 · a risk to the system

Runs happen where the answer key is out of reach

Agents run on a dedicated Linux VM where github.com and GitHub's content hosts resolve to loopback, so an agent under test cannot open, clone or download the benchmark repository during a run. The block is by name: it stops accidental and naive access, not a deliberate bypass, and it says nothing about what a model saw in training.

Results

Every delivery in a table solved the same package, byte for byte — the same SHA-256 — so the numbers compare like for like.

LEB-100-A v1.1 · ai_benchmark.instances.LEB-100-A.name

ai_benchmark.instances.LEB-100-A.desc

mode A · 30 turns edition 2026 evaluated 2026-09-29 matrix 68088abdb7bc…
  1. 1
    Claude Sonnet 5.5 Anthropic · effort xhigh 1 of 3 runs · not official
    825 of 1000 LEB Gold
    • Security 250/250
    • Architecture 50/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 100/100
    • Compatibility 100/100
    • Explanation 46/50

    Penalties: none Discovery 79.2 Brier 0.022 Scorecard

  2. 2
    Claude Fable 5.1 Anthropic · effort xhigh 1 of 3 runs · not official
    781 of 1000 LEB Gold
    • Security 211/250
    • Architecture 25/200
    • Bugs 150/150
    • Performance 150/150
    • Clean code 100/100
    • Compatibility 100/100
    • Explanation 45/50

    Penalties: none Discovery 91.7 Brier 0.020 Scorecard

  3. 3
    Claude Opus 5.5 Anthropic · effort xhigh 1 of 3 runs · not official
    711 of 1000 LEB Silver
    • Security 233/250
    • Architecture 50/200
    • Bugs 139/150
    • Performance 150/150
    • Clean code 25/100
    • Compatibility 70/100
    • Explanation 44/50

    Penalties: none Discovery 91.7 Brier 0.073 Scorecard

  4. 4
    GPT-6-astra OpenAI · effort xhigh 1 of 3 runs · not official
    661 of 1000 LEB Silver
    • Security 224/250
    • Architecture 0/200
    • Bugs 150/150
    • Performance 150/150
    • Clean code 25/100
    • Compatibility 70/100
    • Explanation 42/50

    Penalties: none Discovery 87.5 Brier 0.000 Scorecard

  5. 5
    GPT-5.6-terra OpenAI · effort xhigh 1 of 3 runs · not official
    625 of 1000 LEB Silver
    • Security 211/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 35/50

    Penalties: none Discovery 45.8 Brier 0.000 Scorecard

  6. 6
    GPT-5.6-sol OpenAI · effort xhigh 1 of 3 runs · not official
    612 of 1000 LEB Silver
    • Security 246/250
    • Architecture 0/200
    • Bugs 129/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 32/50

    Penalties: -15 Discovery 70.8 Brier 0.001 Scorecard

  7. 7
    GPT-5.5 OpenAI · effort xhigh 1 of 3 runs · not official
    601 of 1000 LEB Silver
    • Security 198/250
    • Architecture 0/200
    • Bugs 150/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 33/50

    Penalties: none Discovery 70.8 Brier 0.006 Scorecard

  8. 8
    GPT-5.6-luna OpenAI · effort xhigh 1 of 3 runs · not official
    599 of 1000 LEB Bronze
    • Security 207/250
    • Architecture 0/200
    • Bugs 139/150
    • Performance 150/150
    • Clean code 0/100
    • Compatibility 70/100
    • Explanation 33/50

    Penalties: none Discovery 83.3 Brier 0.003 Scorecard

Flaw by flaw

What each agent found and fixed among the planted flaws. The hard ones are flaws of absence — a missing authorization check, a session never regenerated, a file left open on the error path.

  • fixed
  • found, not fixed
  • missed
Flaw Claude Sonnet 5.5 Claude Fable 5.1 Claude Opus 5.5 GPT-6-astra GPT-5.6-terra GPT-5.6-sol GPT-5.5 GPT-5.6-luna
SQL injection in the search SEC-001 · critical · easy 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed 10/10 fixed
Reflected XSS in the search SEC-003 · high · easy 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed
Formula injection in the CSV export SEC-008 · medium · hard 6/6 fixed 2/6 found, not fixed 2/6 found, not fixed 6/6 fixed 0/6 missed 6/6 fixed 0/6 missed 2/6 found, not fixed
Session fixation at login SEC-013 · high · hard 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 5/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed
Unsalted MD5 passwords SEC-014 · high · easy 8/8 fixed 8/8 fixed 8/8 fixed 3/8 found, not fixed 8/8 fixed 8/8 fixed 3/8 found, not fixed 3/8 found, not fixed
Secrets hardcoded in the config SEC-015 · high · easy 8/8 fixed 3/8 found, not fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed
Any ticket readable by id (IDOR) SEC-017 · critical · hard 10/10 fixed 10/10 fixed 10/10 fixed 9/10 fixed 10/10 fixed 9/10 fixed 9/10 fixed 9/10 fixed
Division by zero in the SLA average BUG-001 · high · moderate 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed
File handle leaked on the error path BUG-004 · medium · hard 4/6 fixed 6/6 fixed 5/6 fixed 6/6 fixed 4/6 fixed 4/6 fixed 6/6 fixed 5/6 fixed
One query per ticket for the technician (N+1) PERF-001 · high · moderate 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed 8/8 fixed
A dispatcher that does everything ARCH-002 · high · easy 4/10 found, not fixed 2/10 found, not fixed 4/10 found, not fixed 0/10 missed 0/10 missed 0/10 missed 0/10 missed 0/10 missed
Magic numbers for status and priority ARCH-009 · low · moderate 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed 0/6 missed
Four levels of nested ifs CLN-007 · medium · easy 8/8 fixed 8/8 fixed 2/8 found, not fixed 2/8 found, not fixed 0/8 missed 0/8 missed 0/8 missed 0/8 missed

What stood out

  • Nobody fixed the formula injection in the CSV (SEC-008). Three agents reported it and chose to keep the cells raw for the export's consumers; GPT-5.5 did not report it.
  • Architecture was the weakest category for all four — 25, 50, 0 and 0 of 200. Nobody split the dispatcher that does everything: Fable 5.1 and Opus 5.5 named it and declined to restructure it.
  • Nobody broke the contract mechanically. All four stayed on mysqli, kept the 22 characterization checks green and reported no decoy. Judgement is what separated them: only Fable 5.1 kept compatibility at 100, while each of the other three changed a business value (−30).
  • Passwords and secrets split the field. Fable 5.1 and Opus 5.5 migrated MD5 to password_hash transparently at login; both GPT models left MD5 in place on purpose. Fable 5.1, in turn, kept the secrets in the config as literal fallbacks.
  • GPT-5.5 and GPT-5.6-luna are two points apart, across the Silver/Bronze line — well inside the noise of a single run.

Read this before quoting a number

  • One run per agent. An official LEB score is the median of three independent runs. These are single runs, and a second run can move a total by tens of points.
  • The judge is an AI. Claude Opus 5.5 applied the published rubric to every delivery without knowing which model wrote it — each was anonymised — and the explanation was scored by a separate judge that saw neither the answer key nor the other scores.
  • The judge is also a contestant. Claude Opus 5.5 is one of the agents evaluated, and three of the eight are Claude models — the top three places. Anonymity limits that bias; it does not remove it, since a model can recognise its own style. Every verdict is published with its rationale, flaw by flaw, and the three verdicts changed in review say why; one of them lifts GPT-5.6-terra from last to 5th.
  • The answer key is public. The failure matrix of LEB-100-A has been in the public repository since July 2026. The VM kept it out of reach during the runs, but it may have reached training data: LEB-100-A should be retired for new runs.
  • The benchmark's own tests were fixed. Scoring these runs exposed two defects in the evaluation tooling: an SQL loader that split a statement on a semicolon inside a comment, and CSV checks that read a temporary file the contract never promised. Both were fixed before scoring, the same way for every agent, and are recorded in the repository.
  • Some run parameters were not recorded: the exact model version, the temperature, token counts and cost, and the full logs. Each run marks them as not recorded rather than guessing.

Audit it

Every delivery, mechanical report, verdict and scorecard is in the repository, next to the specification that produced them.