LEB-300 exploratory pilot
An exploratory pilot. This is the median of three runs of one agent, as the protocol asks. The instance itself is still a pilot: its difficulty has not been homologated. Only the aggregate is published: the planted defects, the verdict, the code under test and the answer key stay private.
The first result
-
1
Claude Haiku 4.5 Anthropic · default effort (not configurable) 3 of 3 runs242 of 1000 Failed
- Security 56/250
- Architecture 25/200
- Bugs 29/150
- Performance 0/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 32/50
What it is
LEB is the LLM Engineering Benchmark: an agent is handed a working system with defects planted in it, and is scored on what it finds, explains and fixes without breaking what already worked. LEB-100 is the first level, an application of about 300 lines. LEB-300 is the level above, an application of about 3,000 lines.
Where it stands
The instance is in an exploratory pilot. The result above is the first one published, and other agents follow. Its first runs also served to check the procedure.
What the record does not have
- No checkpoint was taken between the two stages of the task in any of the runs, so it is not on record that the model stayed the same through the first stage. The transcripts show a single model.
- The task text the agent read names the previous scoring-matrix hash. The matrix this run is scored against differs from it only by a header field that marks the instance as active. The scoring is the same.
What is published
This page shows one line per agent: the total, the grade, the score in each category, how many runs it rests on, and the cost and the time. It never shows the list of planted defects, the code under test or the answer key. The same instance has to keep measuring new agents, so those stay private.
Where to read more
The protocol, the scoring and the tools are public in the ai-benchmark repository. The results of the first level are on the LEB-100 page.