“I can’t trust what it gives back.”
Every answer arrives equally confident, right or wrong. SignalPilot audits before anything ships, and shows you the trail and its confidence on every answer.
Open-Source Data-Agent Infrastructure
AI didn’t reduce your work. It moved it: orchestrating harnesses, verifying output. SignalPilot is the infrastructure that makes that easy, and makes the same model more than 2× as correct.
SAME CLAUDE MODEL: 39.5% → 96.9% ACCURATE · #1 ON BOTH PUBLIC BENCHMARKS · EVERY TRANSCRIPT PUBLIC
Every update to your agent stack is evaled before it lands. Nothing breaks quietly.
A Real Example
We asked both for total revenue on a Shopify store’s dbt project.
The $1.3M mistake came from base Claude Code, shipped under its own green check. With SignalPilot the model didn’t get smarter. The process refused to let it guess.
The Frustrations
Every answer arrives equally confident, right or wrong. SignalPilot audits before anything ships, and shows you the trail and its confidence on every answer.
Config sprawl, prompt tuning, rewrites with every model release. SignalPilot ships the blocks: maintained and continuously evaluated so you don’t.
The Building Blocks
The hard part isn’t wiring tools. It’s accuracy. Take all six, or start with the one you’re missing.
The same discipline every run, shipped as a Claude Code or Codex plugin.
Every answer ships with a receipt. Verify your agent’s work at a glance.
Ship updates to your agent stack without worrying about breaking it. Every change is evaled before it lands.
Your agent encodes the tribal knowledge: definitions, quirks, decisions. Every run adds to it.
Agents reach your warehouse through one MCP gateway. Access governed, every query logged.
The same blocks underneath, wherever work happens.
Every block is open source.
Every Way Your Agent Lies
every column checked against your schema before the build
join grain measured before writing, audited after
output compared against the source table
today’s values pulled live before filtering
nothing ships until the audit passes
Continuous Evals
Your warehouse keeps moving. A hand-rolled harness breaks randomly and waits for a human to notice and patch it. SignalPilot pairs continuous auto-evals with a compounding knowledge base: drift gets caught the run it appears, every catch becomes a rule, and accuracy only moves one way.
The Receipts
Same Claude model on both runs. The only difference is the blocks. Every transcript is public.
Same model, with and without the blocks.
#1 on the leaderboard · every transcript public →
The hardest public data-engineering benchmark.
#1 on the live leaderboard · task-by-task results →
Build vs. Buy
Everything we used is published: the workflow, the skills, the verifiers, the eval. Run it against your own harness and see where it breaks. If you’d rather not maintain it, that’s the product.
Get Started
We’re data engineers. Claude Code kept handing us confident, wrong numbers on our own pipelines. So we built the blocks that make it stop.
We host it. Sign up and grab your API key. (Or self-host with docker compose.)
Snowflake, Databricks, Postgres, DuckDB… stored and governed for you.
$ claude plugin marketplace add SignalPilot-Labs/signalpilot-plugin$ claude plugin install signalpilot-dbt@signalpilotEvery query now routes through governance.
$ claude mcp add --transport http signalpilot https://gateway.signalpilot.ai/mcp --header "X-API-KEY: <your-key>"✓ OPEN SOURCE ✓ APACHE 2.0 ✓ SELF-HOST WITH DOCKER COMPOSE