Exmergo logoExmergo
Measured, not claimed

Benchmarks

How Exmergo performs on public analytics-engineering benchmarks. Every run is published under its own dated page, including the baseline it was compared against, so the numbers can be checked rather than taken on trust.

Best result · dbt Labs’ ADE-bench ·

76.0%57 / 75 tasks resolved

Dex + Claude Sonnet 5 resolves more tasks than any other run we measured on dbt Labs’ 75-task ADE-bench.

For context, dbt's published agent skills reported 58% on this benchmark with Opus 4.6.

Newest run · Snowflake’s data-eng-bench ·

57.3% (59 of 103 tasks) on Snowflake’s 103-task data-eng-bench, the top published figure for any Claude Sonnet 5 configuration.

Read the data-eng-bench run

How we benchmark

What does Exmergo publish benchmark results for?
Claims about agent quality are cheap; measured results on a public suite are not. Exmergo runs its products against third-party analytics-engineering benchmarks and publishes every run, including the baseline the run is compared against, so the numbers can be checked rather than taken on trust.
Why is every benchmark run published under its own dated URL?
A benchmark result is only meaningful alongside the date it was measured, the models it used, and the harness configuration. Each run gets a permanent dated page instead of overwriting the previous one, so older results stay citable and improvements can be traced over time.
Which benchmarks does Exmergo run against?
Two suites today: ADE-bench, dbt Labs' 75-task analytics-engineering benchmark, and data-eng-bench, Snowflake's 103-task data-engineering benchmark. Further suites are added as they become relevant to the work Exmergo's agents actually do, and every one of them is a public benchmark published by a third party rather than an internal test set.
Are these results reproducible?
Yes. Each run page lists the agent, the models, the number of attempts per task, and the exact plugin configuration used, along with any prompt the run's image added, reproduced verbatim. The benchmark suites themselves are public, and the raw per-trial output for every run is committed alongside the harness configuration: under experiments/ for ADE-bench and under benchmarks/data-eng-bench for data-eng-bench.