Skip to main content
All resources
Daily playbook
AI-curated · auto-published from public sources

Shadow AI Agents Are Already Running Your Business: A 6-Step Audit

Finance and marketing teams are quietly building AI agents with no IT sign-off. Here's a 6-step audit to find, verify, and govern them before one fails silently.

|7 min read
AI AgentsAI GovernanceSmall Business OperationsAI ReliabilityAI Visibility

As of July 2026, the AI agent building your quarterly deck or answering your customer emails probably wasn't approved by anyone. It was built by someone in finance or marketing over a weekend, and it's still running. That's not a hypothetical — it's the exact problem workflow automation vendor Tines called out this week with the launch of Tines 3B, a product built specifically because "finance, marketing, and more are building dashboards in Claude Code or Codex," often with IT and security completely unaware it's happening.

If you run a local business or a small operation, you don't need a security team's worth of alarm here. You need a Tuesday-afternoon audit. This playbook walks through how to find the agents already touching your business, decide which ones are safe to keep, and put a floor under the ones that aren't.

Why this matters right now

Three things converged in the last 48 hours that make this more urgent than it sounds:

First, the models doing this work are getting harder to reason about. Moonshot AI shipped open weights for Kimi K3, a 2.8-trillion-parameter model now available through Telnyx's inference API — a model so large it needs dedicated GPU infrastructure just to serve, not something a small business runs on a laptop or a $20/month API plan. The bigger and more capable these models get, the less anyone outside a specialized AI shop can meaningfully audit what they're doing inside your workflows.

Second, capability claims are getting harder to trust at face value. A widely discussed benchmark this week (SlopCodeBench) tested Opus 5 specifically to see how it holds up against real-world coding tasks versus marketing claims — the kind of gap AlphaForge has flagged before in why building your own AI stack rarely pencils out the way the sales pitch suggests.

Third, there's a genuinely useful counter-signal from a very different corner of AI: a developer this week published what may be the first formally verified 3D mesh-intersection implementation, built in Lean 4. The entire point of the project was to trust a 93-line mathematical specification instead of roughly 1,000 lines of unverified code generated to do the same job — a better than 10-to-1 ratio of trusted spec to unverified implementation. That's the right instinct for any business owner right now: stop trusting that an agent works because it produced a plausible-looking result, and start checking it against a plain-language spec of what it's actually supposed to do.

The 6-step shadow agent audit

  1. List every tool with an API key or login, not just the ones IT set up. Check billing statements for Anthropic, OpenAI, Zapier, Make, n8n, and any "AI" line item under $50/month — those are the ones nobody asks permission for.
  2. Ask each team what decisions the tool makes without a human clicking approve. Sending an email, updating a CRM record, quoting a price, or replying to a customer review all count. If the answer is "I'm not totally sure," that's the agent to prioritize.
  3. Write a one-paragraph spec for what each agent is supposed to do — before you check what it's actually doing. This mirrors the formally-verified-CSG logic: a short, plain spec you can hold the output against is worth more than a thousand lines of code or prompt you can't fully read.
  4. Check where the output lands. An agent drafting an internal memo carries different risk than one posting to your Google Business Profile, sending invoices, or answering a prospective customer directly. Rank each agent by how public and how reversible its output is.
  5. Set a spend and model ceiling for each agent. With frontier models now running into the trillions of parameters and dedicated-infrastructure territory, an agent that silently upgrades itself to a bigger, pricier model can quietly change your monthly bill. Cap it, or route it through a gateway that enforces a ceiling automatically.
  6. Assign one human owner per agent, by name. Not "marketing" — a person. If nobody can tell you who owns an agent within thirty seconds, that agent gets paused until someone claims it.

Quick audit checklist

  • Every AI tool with billing access is listed in one place, not scattered across personal cards
  • Every agent that touches a customer or a public page has a named human owner
  • Every agent has a one-paragraph spec you could hand to someone else and have them check the output against it
  • Every agent has a spend ceiling, not an open-ended API key
  • Someone reviews agent output at least weekly — not just when a customer complains

Common pitfalls

The most common mistake is treating this as an IT problem to delegate and forget. The Tines data point this week is specific: these agents are being built by finance and marketing, not by IT, which means IT can't audit what it doesn't know exists. The audit has to start with a conversation across departments, not a ticket in a queue.

The second mistake is assuming a bigger, newer model is automatically safer or more accurate. Benchmarks like SlopCodeBench exist precisely because model marketing and model reliability are two different things — a 2.8-trillion-parameter model is impressive, not inherently correct.

The third mistake is skipping the spec step because writing it down feels slower than just watching the output. It isn't. A short spec is the cheapest verification tool you have, and it's the difference between catching a bad agent decision in a Tuesday review versus a customer catching it for you.

None of this requires becoming an AI engineer. It requires knowing which of your customer-facing surfaces are already being touched by an agent nobody signed off on — and whether the version of your business showing up when a customer asks an AI assistant for the best option nearby is accurate, current, and something you'd actually approve.

If you're not sure what that looks like right now, start with the free 24-hour free AI Visibility Report — it'll show you exactly what AI assistants are already saying about your business, unaudited or not.

Want this built for you?

Pick a tier, pick an agent. Live in 48 hours.