Wendl
AI Bill Teardown
You are paying frontier-model prices for traffic that mostly doesn't need them.
Today / mo
$1,625
all Opus, every prompt
With Wendl / mo
$553
routed per prompt
You save
$1,07366%
≈ $12,870 / year
Quality held at 4.4/5 vs 4.6/5 for all-Opus, judged on your own prompts.
The proof — cost × quality on your traffic
Cheaper is easy; cheaper at the same quality is the point. Every option below is graded on your own prompts. Anything on the dashed frontier is an efficient choice — nothing beats it on both axes. Wendl sits on it, near the top on quality at a fraction of the cost.
Where your traffic actually goes
Wendl reads each prompt and sends it to the cheapest tier that can handle it. On this sample, 25% ran free on local models and only 15% needed a frontier model.
Local · small 15% llama3.1:8b
Local · mid 10% llama3.1:8b
Cheap cloud 25% claude-haiku-4-5-20251001
Mid cloud 35% claude-sonnet-5
Deep cloud 15% claude-opus-4-8
What we'd change
- You send 100% of traffic to Opus today. On this sample, only 15% of prompts actually need a frontier model — the rest are over-served.
- 25% of your prompts are routine enough to run on a local model at $0 marginal cost — the single biggest lever in your bill.
- Quality holds: 4.4/5 routed vs 4.6/5 all-Opus on your prompts — a 0.2-point gap for a 66% cost cut.
- Set a hard monthly budget and let low-priority traffic (cron, triage, acks) default to the cheapest tier — spend follows priority, not habit.
Want this run on your real traffic?
We set up the routing, prove the numbers on your data, and keep it tuned as models change.
How we measured this
- Measured on 200 of your prompts (acme-sample.jsonl), projected to 50,000 prompts/month.
- Quality graded 1–5 by an independent judge model (claude-sonnet-5); higher is better.
- Illustrative figures for a fictional team — a real teardown uses your own traffic and prices.
- Prices are current provider list rates (refreshed from OpenRouter). Token counts estimated from your prompts; a live run on your traffic tightens them.
- “Local” = an open-weight model on your own hardware: $0 marginal cost, data never leaves your box.