Benchmarks
How Exmergo performs on public analytics-engineering benchmarks. Every run is published under its own dated page, including the baseline it was compared against, so the numbers can be checked rather than taken on trust.
Best result · dbt Labs’ ADE-bench ·
Dex + Claude Sonnet 5 resolves more tasks than any other run we measured on dbt Labs’ 75-task ADE-bench.
For context, dbt's published agent skills reported 58% on this benchmark with Opus 4.6.
Newest run · Snowflake’s data-eng-bench ·
57.3% (59 of 103 tasks) on Snowflake’s 103-task data-eng-bench, the top published figure for any Claude Sonnet 5 configuration.
Read the data-eng-bench run→Benchmarks we run against
Public suites published by third parties, not internal test sets.
ADE-bench
Analytics Engineering Benchmark, by dbt Labs
dbt Labs' 75-task analytics-engineering benchmark, scored on whether the dbt project's tests actually pass.
75 tasks · 1 published run
About dbt Labs’ ADE-benchdata-eng-bench
Data Engineering Agent Benchmark, by Snowflake
Snowflake's 103-task data-engineering benchmark, scored pass/fail by whether the dbt project's own pytest suite runs clean.
103 tasks · 1 published run
About Snowflake’s data-eng-benchHow we benchmark
- What does Exmergo publish benchmark results for?
- Claims about agent quality are cheap; measured results on a public suite are not. Exmergo runs its products against third-party analytics-engineering benchmarks and publishes every run, including the baseline the run is compared against, so the numbers can be checked rather than taken on trust.
- Why is every benchmark run published under its own dated URL?
- A benchmark result is only meaningful alongside the date it was measured, the models it used, and the harness configuration. Each run gets a permanent dated page instead of overwriting the previous one, so older results stay citable and improvements can be traced over time.
- Which benchmarks does Exmergo run against?
- Two suites today: ADE-bench, dbt Labs' 75-task analytics-engineering benchmark, and data-eng-bench, Snowflake's 103-task data-engineering benchmark. Further suites are added as they become relevant to the work Exmergo's agents actually do, and every one of them is a public benchmark published by a third party rather than an internal test set.
- Are these results reproducible?
- Yes. Each run page lists the agent, the models, the number of attempts per task, and the exact plugin configuration used, along with any prompt the run's image added, reproduced verbatim. The benchmark suites themselves are public, and the raw per-trial output for every run is committed alongside the harness configuration: under experiments/ for ADE-bench and under benchmarks/data-eng-bench for data-eng-bench.
