On dbt Labs' 75-task analytics-engineering benchmark, Dex + Claude Sonnet 5 resolves more tasks than any other run we measured, and beats dbt's own published agent skills.
Best run: Claude Sonnet 5 with dex
Read the run→Published by dbt Labs
Analytics Engineering Benchmark
dbt Labs' 75-task analytics-engineering benchmark, scored on whether the dbt project's tests actually pass.
ADE-bench hands an agent a real dbt project running on DuckDB and asks it to fix a broken model, build a new one, or extend the semantic layer. Each of the 75 tasks is graded by running the project's own tests afterwards, so the benchmark measures completed engineering work rather than plausible-looking SQL. Tasks span eight domains, and there is no partial credit: a task counts only when the resulting project genuinely builds and passes.
Latest result ·
Dex + Sonnet 5 leads at 76%, for about 2.5x less than Fable 5 and ~17% less than Opus 4.8. With Dex, accuracy holds in a 72–76% band across all three models while cost ranges from $36 to $92, so the practical call is to run an inexpensive model.
For context, dbt's published agent skills reported 58% on this benchmark with Opus 4.6.
Read the full 21 July 2026 run→Every ADE-bench run Exmergo has published, newest first. Each run keeps its own permanent page, so earlier results stay citable.
Results measured by Exmergo on dbt Labs’ ADE-bench.