Same model, with and without SignalPilot.
Both runs: Claude Code with Sonnet 4.6 · every transcript public →
We build the evals once from your stack, then run every candidate through them: Claude, GPT, Gemini, warehouse agents, BI agents. You get pass rates per question, evidence for every answer, and a decision you can defend.
24 graded questions · 4 candidates · same context, same sandbox
Every cell links to the transcript and the evidence graph. The context and the evals stay in your repo, whichever candidate you pick.
Why a bakeoff needs evals
Our public results, for reference. Every team we talk to asks the second question, and only your own graded questions answer it.
#1
on both public benchmarks
Spider 2.0-DBT and ADE-Bench
96.9%
ADE-Bench, tasks passed
Claude Code alone: 39.5%, same model
75%
Spider 2.0-DBT, tasks passed
next best 60.29%
Same model, with and without SignalPilot.
Both runs: Claude Code with Sonnet 4.6 · every transcript public →
The hardest public data-engineering benchmark.
*75% submitted in August; official leaderboard shows 65.6% (May).
What you get
Model routing
The report gives you pass rates per task group and cost per task for every candidate. Route each group to the cheapest model that clears the bar and the expensive model only runs where it earns its keep.
| Model | Month-end close24 tasks | Pipeline review18 tasks | Refunds12 tasks | Retention10 tasks | Failure cases8 tasks | Cost / task |
|---|---|---|---|---|---|---|
| Claude Opus | 100% | 96% | 100% | 100% | 92% | $1.20 |
| GPT-5.6 | 92% | 92% | 96% | 96% | 79% | $0.60 |
| Claude Sonnet | 96% | 92% | 100% | 100% | 83% | $0.40 |
| GLM 5.3 | 83% | 88% | 96% | 92% | 58% | $0.12 |
| Gemini Flash | 75% | 79% | 96% | 92% | 42% | $0.08 |
| DeepSeek | 79% | 83% | 92% | 90% | 50% | $0.06 |
✓ marks the cheapest model that clears the bar for that group. Where no cheaper model clears it, the top model keeps the slot.
$1.20 / task
98.1% of tasks pass · $86.40 a month
$0.39 / task
93.1% of tasks pass · $27.72 a month
68%
of model spend, with every task group above the bar. Rerun when the next model ships.
Illustrative figures for a fictional project. Your bakeoff produces your own table, on your questions, with your token prices.
Every candidate answers the same graded questions in the same sandbox with the same context, so the pass rates are comparable and the cost per task is real.
You set the pass rate a task group must clear. The cheapest model above it takes the slot. Failure cases usually stay with the strongest model; routine groups rarely need it.
The evals and the context stay in your repo. When a new model lands, rerun the grid and move the slots that changed.
About two weeks from access to report for a scoped set of questions.
The evals and the context layer are yours. Rerun the bakeoff when the next model ships.