LEB-100-A v1.1 · ai_benchmark.instances.LEB-100-A.name
ai_benchmark.instances.LEB-100-A.desc
-
1
Claude Sonnet 5.5 Anthropic · effort xhigh 1 of 3 runs · not official825 of 1000 LEB Gold
- Security 250/250
- Architecture 50/200
- Bugs 129/150
- Performance 150/150
- Clean code 100/100
- Compatibility 100/100
- Explanation 46/50
-
2
Claude Fable 5.1 Anthropic · effort xhigh 1 of 3 runs · not official781 of 1000 LEB Gold
- Security 211/250
- Architecture 25/200
- Bugs 150/150
- Performance 150/150
- Clean code 100/100
- Compatibility 100/100
- Explanation 45/50
-
3
Claude Opus 5.5 Anthropic · effort xhigh 1 of 3 runs · not official711 of 1000 LEB Silver
- Security 233/250
- Architecture 50/200
- Bugs 139/150
- Performance 150/150
- Clean code 25/100
- Compatibility 70/100
- Explanation 44/50
-
4
GPT-6.1-sol OpenAI · effort xhigh 1 of 3 runs · not official666 of 1000 LEB Silver
- Security 228/250
- Architecture 25/200
- Bugs 150/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 43/50
-
5
GPT-6-astra OpenAI · effort xhigh 1 of 3 runs · not official661 of 1000 LEB Silver
- Security 224/250
- Architecture 0/200
- Bugs 150/150
- Performance 150/150
- Clean code 25/100
- Compatibility 70/100
- Explanation 42/50
-
6
GPT-5.6-terra OpenAI · effort xhigh 1 of 3 runs · not official625 of 1000 LEB Silver
- Security 211/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 35/50
-
7
GPT-5.6-sol OpenAI · effort xhigh 1 of 3 runs · not official612 of 1000 LEB Silver
- Security 246/250
- Architecture 0/200
- Bugs 129/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 32/50
-
8
GPT-5.5 OpenAI · effort xhigh 1 of 3 runs · not official601 of 1000 LEB Silver
- Security 198/250
- Architecture 0/200
- Bugs 150/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 33/50
-
9
GPT-5.6-luna OpenAI · effort xhigh 1 of 3 runs · not official599 of 1000 LEB Bronze
- Security 207/250
- Architecture 0/200
- Bugs 139/150
- Performance 150/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 33/50
-
10
MiniMax-M3 MiniMax · default effort (not configurable) 1 of 3 runs · not official460 of 1000 LEB Bronze
- Security 185/250
- Architecture 0/200
- Bugs 129/150
- Performance 56/150
- Clean code 0/100
- Compatibility 70/100
- Explanation 20/50
Flaw by flaw
What each agent found and fixed among the planted flaws. The hard ones are flaws of absence — a missing authorization check, a session never regenerated, a file left open on the error path.
- fixed
- found, not fixed
- missed
| Flaw | Claude Sonnet 5.5 | Claude Fable 5.1 | Claude Opus 5.5 | GPT-6.1-sol | GPT-6-astra | GPT-5.6-terra | GPT-5.6-sol | GPT-5.5 | GPT-5.6-luna | MiniMax-M3 |
|---|---|---|---|---|---|---|---|---|---|---|
| SQL injection in the search | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed | 10/10 fixed |
| Reflected XSS in the search | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 6/8 found, not fixed |
| Formula injection in the CSV export | 6/6 fixed | 2/6 found, not fixed | 2/6 found, not fixed | 2/6 found, not fixed | 6/6 fixed | 0/6 missed | 6/6 fixed | 0/6 missed | 2/6 found, not fixed | 0/6 missed |
| Session fixation at login | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 5/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed |
| Unsalted MD5 passwords | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 3/8 found, not fixed | 8/8 fixed | 8/8 fixed | 3/8 found, not fixed | 3/8 found, not fixed | 3/8 found, not fixed |
| Secrets hardcoded in the config | 8/8 fixed | 3/8 found, not fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 6/8 found, not fixed |
| Any ticket readable by id (IDOR) | 10/10 fixed | 10/10 fixed | 10/10 fixed | 9/10 fixed | 9/10 fixed | 10/10 fixed | 9/10 fixed | 9/10 fixed | 9/10 fixed | 10/10 fixed |
| Division by zero in the SLA average | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed |
| File handle leaked on the error path | 4/6 fixed | 6/6 fixed | 5/6 fixed | 6/6 fixed | 6/6 fixed | 4/6 fixed | 4/6 fixed | 6/6 fixed | 5/6 fixed | 4/6 fixed |
| One query per ticket for the technician (N+1) | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 8/8 fixed | 3/8 found, not fixed |
| A dispatcher that does everything | 4/10 found, not fixed | 2/10 found, not fixed | 4/10 found, not fixed | 2/10 found, not fixed | 0/10 missed | 0/10 missed | 0/10 missed | 0/10 missed | 0/10 missed | 0/10 missed |
| Magic numbers for status and priority | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed | 0/6 missed |
| Four levels of nested ifs | 8/8 fixed | 8/8 fixed | 2/8 found, not fixed | 0/8 missed | 2/8 found, not fixed | 0/8 missed | 0/8 missed | 0/8 missed | 0/8 missed | 0/8 missed |
What stood out
- Nobody fixed the formula injection in the CSV (SEC-008). Three agents reported it and chose to keep the cells raw for the export's consumers; GPT-5.5 did not report it.
- Architecture was the weakest category for all four — 25, 50, 0 and 0 of 200. Nobody split the dispatcher that does everything: Fable 5.1 and Opus 5.5 named it and declined to restructure it.
- Nobody broke the contract mechanically. All four stayed on mysqli, kept the 22 characterization checks green and reported no decoy. Judgement is what separated them: only Fable 5.1 kept compatibility at 100, while each of the other three changed a business value (−30).
- Passwords and secrets split the field. Fable 5.1 and Opus 5.5 migrated MD5 to
password_hashtransparently at login; both GPT models left MD5 in place on purpose. Fable 5.1, in turn, kept the secrets in the config as literal fallbacks. - GPT-5.5 and GPT-5.6-luna are two points apart, across the Silver/Bronze line — well inside the noise of a single run.