As of July 2026, two stories moving through the developer world this week point at the same blind spot for anyone running, or planning to run, an AI agent stack: the overhead you do not see is often bigger than the work you do see, and the fix is not a smarter model, it is a smarter audit of what is running underneath it.
The 4.7x tax nobody bills you for
A team at Systima ran a controlled comparison between two popular coding agents, Claude Code and OpenCode, logging every token that moved between the tool and the model before any real work happened. Their finding: Claude Code sent roughly 33,000 tokens of scaffolding — system prompts, tool definitions, context loading — before it ever read the developer's actual request. OpenCode did the same job in about 7,000 tokens. That is a 4.7x overhead tax paid before a single line of real work starts, and it compounds on every call, all day, every day, regardless of how simple or complex the actual request was.
Nobody notices this in the moment. The invoice arrives at the end of the month as "usage was higher than expected," and most teams tune the prompt, not the plumbing underneath it. But the study makes a point local business owners running AI agent tools for booking, content, or outreach need to hear directly: the tool wrapped around the model can cost you more than the model itself, and two tools doing the identical job can differ by nearly 5x before either one produces a result. We have made this exact case before — see the hidden cost driving up your AI agent's bill — and this week's data is a clean, measured confirmation of it.
Same job, 2.2x faster, 27% cheaper — just from switching engines
The second story comes from Ploy.ai, which migrated a production customer-facing agent from its prior model to GPT-5.6. The result, on the same workload, same task definitions, same customers: 2.2x faster response times and a 27 percent drop in cost. Nothing about what the agent does for the business changed — no new features, no new prompts. What changed was the engine underneath it, and the business kept the savings.
Put those two data points together and a pattern shows up: neither the overhead tax nor the win came from the AI layer everyone talks about in sales pitches. Both came from infrastructure decisions — which scaffolding wraps the model, which model runs the workload — decisions that typically get made once, at launch, and then get ignored for months while the bill quietly reflects them either way.
What this means if you're running, or buying, an agent stack
If you are a business owner with an AI agent handling your booking, review responses, or outreach, you probably do not see token counts. You see a subscription price or a monthly invoice from whoever built it. That is exactly why this matters: the 4.7x gap and the 27 percent swing are invisible from where you sit, but they show up in your bill or your vendor's margin regardless. Ask whoever runs your stack two questions this week:
- What is our overhead ratio — tokens spent on scaffolding versus tokens spent on the actual customer request?
- When did we last benchmark the model underneath against a newer, faster, or cheaper release?
If nobody can answer either one with a number, you are paying a tax you cannot see and have not asked anyone to measure.
This is also the argument for hiring a specialist over duct taping a stack together from tutorials and forum threads. An agent build that nobody re-benchmarks against faster, cheaper model releases does not stay competitive — it just gets more expensive relative to the alternative, quietly, for as long as nobody checks the math. The fix is not complicated: pick your scaffolding for overhead, not just feature lists, and put re-checking the model underneath on a recurring calendar item instead of treating it as a one-time decision made at launch.
What this means if you're weighing AI marketing or an agent build: the tool wrapping the model matters as much as the model itself, and that holds whether the agent is writing code or answering your customers, so get someone to measure your actual overhead before you assume the invoice reflects the work being done.
Not sure whether your business is even showing up when customers ask ChatGPT, Claude, or Perplexity for the best option nearby? Get a free AI Visibility Report and find out in 24 hours.