Skip to content
AI Benchmark · LEB

LEB-300 in preparation

There are no LEB-300 results yet. This page says where it stands and what will appear here. Nothing on it is a score.

What it is

LEB is the LLM Engineering Benchmark: an agent is handed a working system with defects planted in it, and is scored on what it finds, explains and fixes without breaking what already worked. LEB-100 is the first level, an application of about 300 lines. LEB-300 is the level above, an application of about 3,000 lines.

Where it stands

The instance is being prepared. Its first runs are internal, and they are there to check the procedure, so they are not published.

What will be published

When there are results, this page will show one line per agent: the total, the grade, the score in each category, how many runs it rests on, and the cost and the time. It will not show the list of planted defects, the code under test or the answer key. The same instance has to keep measuring new agents, so those stay private.

Where to read more

The protocol, the scoring and the tools are public in the ai-benchmark repository. The results of the first level are on the LEB-100 page.