Skip to content
AI Benchmark · LEB

LEB-300 exploratory pilot

An exploratory pilot. This is the median of three runs of one agent, as the protocol asks. The instance itself is still a pilot: its difficulty has not been homologated. Only the aggregate is published: the planted defects, the verdict, the code under test and the answer key stay private.

The first result

  1. 1
    Claude Haiku 4.5 Anthropic · default effort (not configurable) 3 of 3 runs
    242 of 1000 Failed
    • Security 56/250
    • Architecture 25/200
    • Bugs 29/150
    • Performance 0/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 32/50

    Cost US$ 0.96 a run Session 13.4 min

What it is

LEB is the LLM Engineering Benchmark: an agent is handed a working system with defects planted in it, and is scored on what it finds, explains and fixes without breaking what already worked. LEB-100 is the first level, an application of about 300 lines. LEB-300 is the level above, an application of about 3,000 lines.

Where it stands

The instance is in an exploratory pilot. The result above is the first one published, and other agents follow. Its first runs also served to check the procedure.

What the record does not have

  • No checkpoint was taken between the two stages of the task in any of the runs, so it is not on record that the model stayed the same through the first stage. The transcripts show a single model.
  • The task text the agent read names the previous scoring-matrix hash. The matrix this run is scored against differs from it only by a header field that marks the instance as active. The scoring is the same.

What is published

This page shows one line per agent: the total, the grade, the score in each category, how many runs it rests on, and the cost and the time. It never shows the list of planted defects, the code under test or the answer key. The same instance has to keep measuring new agents, so those stay private.

Where to read more

The protocol, the scoring and the tools are public in the ai-benchmark repository. The results of the first level are on the LEB-100 page.