Mid-Market AI

What is the AI Control Plane? And Why Yours Should Be Federated | Mid-Market AI | Episode 111

Paragon Season 2 Episode 111

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 32:40

The control plane is the token cost controller is AI sovereignty is AI security is the ROI unlock. It's all the same conversation.

I called the SaaSpocalypse hype. Then in a single week, several clients and partners told me they're building custom CRMs or already run one, and on July 1 Gartner priced the trend at $234 billion in enterprise software spend at risk. When software goes headless and agentic, the interface stops being where you govern it. Every vendor, from the hyperscalers to a crop of startups, wants to sell you the replacement as a single product. This episode is our approach instead: the Federated Control Plane, eight independent control domains built out of the evidence.

Covered: the Cloud Security Alliance and Singapore IMDA definitions of the agentic control plane, the loop-bill-drift failure pattern, agent token economics (4x to 15x chat, 100x cost for 1% accuracy), 84% prompt injection success rates, Palantir's sovereignty paper and control liquidity, the vendor land grab from Microsoft Agent 365 to ServiceNow AI Control Tower, who watches the watchers, the eight domains, the reversibility test, and the PE angle: portfolio governance, agent-washing diligence, and sponsor-level EU AI Act exposure.

For portco CEOs, PE operating partners, CFOs staring at unexplainable AI bills, and operators deciding which control plane to trust. Answer: several, on purpose.

Mid-Market AI is produced by Paragon Technology Solutions.

Paragon - Managed Intelligence Provider (MIP™)


SPEAKER_00

When the word SASPOCalypse started going around, I called it hype. Enterprise software has survived every death sentence handed to it. Last week though, changed my mind. Several clients and partners within a few days told me that they're building a custom CRM or already are running one. Now these are lower mid-market companies. Salesforce is probably not going anywhere anytime soon, but if you're in a niche industry, you could now have the option of using AI to code what your business or customers specifically need instead of shoehorning enterprise or SaaS into shape through years of customization. Then on July 1st, Gartner put a price tag on it. $234 billion of enterprise software spending at risk by 2030 from what they're calling agentic arbitrage. Agents completing tasks across systems and bypassing the interfaces entirely. We've covered headless AI and the SaaSpocalypse in prior episodes. Our clients and partners just got there before the analysts did. What the number doesn't price in is when the software goes headless and agentic, and some of it is the software you wrote last quarter or last week, the interface stops being where you govern it. Something else has to. And the custom built stuff arrives with no governance attached at all. Every vendor in the market and every hyperscaler and every package is going to want to sell you their version of that something as a single product or wrapped in. Today I'm giving you our recommended approach instead. We call it the Federated Control Plane. And by the end of this episode, you should be able to sketch it on a whiteboard. Welcome. You're listening to Mid-Market AI. I'm Ariel Jalali, CEO of Paragon. Our managed intelligence provider, or MIP, is a managed AI solution that puts forward-deployed chief AI officer-led data and AI engineering pods inside of PE backed and mid-market companies. We've covered in prior episodes the harness, headless AI. Our last episode was about AI sovereignty, tremendously important and will be growing in importance over the next few months. Today is the fourth leg, the control plane, and it's the framework those episodes were building towards. So we get the call when AI pilots go sideways. Bad signal. Help. They're burning too much money in AI tokens. They're missing the objectives that they were deployed for. Maybe they're drifting in accuracy while nobody's measuring and they're not hitting their ROI targets. The client thinks that they have four problems: a cost problem, a security problem, a sovereignty problem, and an ROI problem. In fact, they have one. It's all the same conversation. I wrote a single sentence in my notes that says exactly how I feel about this, and I'm saving it for the close. So every vendor in the market will soon be selling you a control plane, which is exactly why no single one can be yours. You're getting a federal control plane whether you want one or not. The only decision is whether it's designed or accidental. First, let's get some definitions together because the term got real ones this year, one industry definition and one government definition, and they converge. The Cloud Security Alliance, the nonprofit that writes cloud security standards, made securing the agentic control plane its mission for 2026 and defined it as the layer that governs how autonomous agents exist and operate. Identity, authorization, orchestration, runtime behavior, and trust. Then there's Singapore. In January at Davos, Singapore's Infocom Media Development Authority, or IMDA launched the Model AI Governance Framework for Agentic AI. The world's first government framework, written specifically for AI agents, updated in May after feedback from over 60 organizations. Singapore never says the words control plane. Instead, they define what every agent must have: access controls, guardrails, human approvals, logging, monitoring, and a verifiable identity, an auto trail showing which agent acted under whose authorization, with action boundaries and checkpoints where a human must approve. The framework is voluntary, but you remain legally accountable for what your agents do. And in order to do business with the Singapore government, you have to comply with this. A government just wrote the control plane's job description and attached your name to the outcome. So the layer is real, your agents and their harnesses are the data plane, they do the work, and the control plane decides who may act, what it costs, whether it's allowed, whether it's safe, and what gets recorded. The harness is per workflow, which we've covered in prior episodes. The control plane is per company, one layer for every harness, every agent, every vendor. How does it meter tokens? Every model call passes through a gateway that knows who called, what it cost, which budget it draws against. No gateway, no meter. And you find out what the quarter costs when the invoice arrives and your CFO is losing their minds. You can have 10 great harnesses and still have no idea what Tuesday costs you. Both definitions share one requirement. The decision happens before the action. Singapore's checkpoints sit in front of the agent's next step. A budget check after the tokens are spent is an invoice. A safety check after the tool call executes is a postmortem. Control means before. One thing both definitions quietly assume the control plane is one thing, the centralized controller. The episode shows why it can't be, and why you should build something else instead. Eight independent domains, each with its own authority. Listen while I call them out. The call we get has three recurring symptoms, usually, and each with real numbers behind it, each one maps to a missing domain. There's the loop, the bill, the drift. Those are the three. Let's talk about the loop. There's a documented GitHub issue on the Open Hands project where an agent got stuck apologizing. Eleven plus retry is two cents each. A quarter burned in 36 seconds before someone killed it. At production scale, a CIO analysis this spring priced three-hour recursive agent loop at about $3,700, or $37,000 per incident at 10 agents. A loop is what happens when nothing holds budget authority over a running agent. In the Federation, that authority is an economic disaster. Let's talk about the bill. The Wall Street Journal reported companies hitting their annual budget in three months, and Microsoft reportedly canceled most of its internal cloud code licenses six months after rollout, partially because of cost. Goldman Sachs projects token consumption grows 24x by 2030. That same CIO analysis found bills running 40% over forecast or more, one institution at 3x. A bill nobody can explain means nobody can answer what happened and why. And that's the observability domain. Let's talk about the drift. A study that tracked 14 agent models over 18 months found only small reliability improvements and open-ended tasks, barely any. And the way most teams evaluate running a task once and checking if it passed overestimates real reliability by 20 to 40 percent. The pilot that impressed everyone in week one is a different system by month six. And without evaluation infrastructure, nobody notices until the outcomes do. The question that everybody missed was how much confidence does this output deserve? And it belongs to the trust domain. These are three symptoms in three domains. We've got five to go. So in terms of token cost conversation and accuracy conversation, they're kind of the same conversation. Let's start with the cost because the numbers are measured and not modeled. Anthropic's own engineering team publish production data from their research system. A single agent uses about four times the tokens of a chat interaction. Multi-agent systems use about 15 times. The company that gets paid by token is telling you its best product multiplies your token bill. And Palantir built a whole paper on AI sovereignty attacking that incentive. And we we actually agree with a lot of its principles wholeheartedly. Princeton built the holistic agent leaderboard. Nearly 22,000 agents run across nine models and nine benchmarks, $40,000 of their own money. The headline found that agents can be a hundred times more expensive while being 1% better with a 400x cost spread on identical coding tasks depending on the model and the scaffold. Last episode I said cheap per token, expensive per outcome. Princeton just put a leaderboard behind this idea. The savings are just as measured, and the freshest numbers are from last month. GitHub published results from Copilot's production fleet in June. A 94% cash hit rate, cash tokens up to 10 times cheaper, a router cutting the costs over 3x at equal accuracy by sending each task to the cheapest model that can handle it. This routing concept is going to come up over and over, and it's something that we're putting into our reference architectures at our client installations. And 43 to 62% token reductions per workflow. The question of which model should handle a task is its own authority in the Federation, the routing domain, and it cannot belong to the company selling you the tokens. You don't let the utility run its own meter. Accuracy belongs in the same beat because the unreliability factor is a cost multiplier. Every retry is more tokens. Meter, the organization, found frontier models succeed nearly 100% of the time on tasks a human finishes in under four minutes, and under 10% on tasks beyond four hours. This is the long window. Andre Karpathi calls it the March of Nines. At 99% per step reliability, a hundred step workflow succeeds 36% of the time. That was a mouthful, so let's let's come at it again. At 99% per step reliability, a hundred-step workflow succeeds only 36% of the time. That's crazy. And per step error isn't even constant, it rises because models condition on their own earlier mistakes. The countermeasures exist. We have some verification steps, some confidence scoring, escalation to humans in the loop when confidence is low. We have some continuous evals watching for drift. That's the trust domain from the last bit that we talked about. Now with its full job to score the confidence, decide when a human reviews, catching the drift before the outcomes do. So we've talked about four domains so far economics, observability, trust, and routing. The next beat forces three more. The average prompt injection attack success rate in agent security bench pre-reviewed at ICLR is 84%. Someone plants instructions in your context, your agent reads, and 84% of the time the agent follows the attacker instead of you. The defenses don't really hold up either. Researchers across OpenAI, Anthropic, and Google DeepMind jointly tested 12 published defenses that had reported near zero attack success. Under adaptive attacks, most broke above 90%, and a 500-person human red team got to 100. The exploits are already public. A single poison GitHub issue caused an agent to read private repository and post a Stripe API key publicly within seconds. No alert to the user. Echo leaked, the first zero-click agent exfiltration CVE, needed one crafted email to make Microsoft 365 Copilot handover OneDrive and SharePoint data with no user interaction. Simon Williamson calls the exposure condition the lethal trifecta. And we've talked about this before. When you have private data access, exposure to untrusted content, and the ability to communicate it externally, the agent becomes really dangerous. That's the trifecta. Any two of these is manageable, all three, and you've got one crafted email from a breach. Most deployed agents have all three on day one. The research forces one conclusion. The model cannot be trust trusted to defend itself. So safety lives outside the model in three authorities. Who is this agent? What scope does it hold? Identity, what data may it see and change, data governance. Is this specific action safe? Watching for injection, exfiltration, poison tools. This is security. The Singapore standards agree, putting access controls and guardrails first on their list. Seven domains now, one to go, and it comes from the lawyers. Sovereignty was in our last episode, so I'll add only what's new. Palantir published the strongest statement of institutional sovereignty anyone has ever put on paper. Institutional sovereignty in the age of AI, and I'm on the side of this paper wholeheartedly. Their line is sovereignty is your alpha. Their core warning is that model providers have a structural incentive to migrate your operational know-how into their weights, where it can be leased to your competitors or used to replace you. Their tell is the pricing. If the provider's incentives were aligned with yours, why do they charge for a token instead of a share of the value they created? Remember the 15x multiplier and asks who benefits. The prescriptions read as a checklist in the paper. Treat any frontier model without zero data retention agreement as extraction prone. Every lab, anthropic, open AI, Google, keep model liquidity so you can switch providers and negotiate from strength. Prefers structural assurance compute that you own or can verify over contractual promises and run every workload through a control layer the institution owns. Model agnostic routing, granular permissions, audit logs, reversible branching, those are domains you've already met. The most sovereignty committed vendor in the industry independently derived half of our federation. The eighth domain arrives with the rules. No customer data leaves the region, no agent approves payments above a threshold, high risk output requires human review. Whether an action should be allowed as its own authority, the policy domain, where Singapore's bounded autonomy lives, and so does the EU AI Act. That's all eight. One test before ROI, and I'll use it for the rest of the episode. Can you reverse each decision that you make? We call these two-door decisions, and we're big fans of them in our recommendations to our clients. Every one of Palantir's prescriptions is a machine for making AI decisions reversible. Model liquidity makes the model choice reversible. Branching makes the agent actions reversible. Structural assurance means the exit exists physically, whatever the contract says. Sovereignty and speed are the same purchase. The company that can reverse any AI decision is the company that can afford to make decisions fast. And note the first mover on identic governance was a sovereign state, not a vendor. Now ROI. Straight from my CFOs out there. The Fed CFO survey in December, 77% of large firms invested in AI, median firm reporting no change in outcomes. Deloitte says 21% of finance chiefs using AI report clear measurable value. 400 CFO surveys surveyed half will cut AI investment that can't show ROI within a year. Even Jamie Diamond says JP Morgan spends about $2 billion a year on AI and gets about $2 billion back and calls for the net present value almost worthless. The measurement problem is actually structural. You cannot unlock ROI that you cannot see. COSO, COSO, the internal controls body that your auditors answer to, now tells finance teams to demand logging and traceability of prompts, inputs, outputs, and approvals, per decision cost accounting, named cost owners. That's the observability domain doing the work of the ROI unlock, and the unlock is the ledger. So eight domains forced into existence by the evidence cost forced economics and routing, drift forced trust, attacks forced identity, data governance and security, the regulators forced policy, and the CFO and auditors forced observability. This list wasn't designed by us at a whiteboard, the failures did the job, which brings us to the vendors because they've noticed the same layer. Everyone is selling you a control plane, which is why you need a federated one. CRN's Agentic AI list this month put it in four words. Control is now the product. The sellers by tier. First, the hyperscalers. Microsoft launched Agent 365 in November, and their tagline is word for word, a control plane for AI agents. AWS shipped Bedrock Agent Core. Google is consolidating its portfolio into the Gemini Enterprise Agent Platform. Now the SaaS platforms. Salesforce embeds one and Agent Force. ServiceNow actually made the boldest claim in the market. At Knowledge 2026, a couple months ago, they positioned their AI control tower as a universal control plane. Every agent from every end, every vendor reporting it to one tower. And it bought three companies to build it. They even convinced Microsoft to route Agent 365 metering through it. Now the Enterprise Suites. IBM announced the opposite bet on the same day from a different stage. Keep your existing tools. IBM Concert correlates all of them. Oracle, SAP, and UiPath are in the same bucket. Two giants opposing control architectures the same day in May, and both still ask you to centralize on the announcing vendor. The frameworks and the labs. Langchain positions Langsmith as the control plane. OpenAI ships the agents SDK. The model providers are becoming application infrastructure, which means that the company selling you the tokens is also now offering you how to govern how to spend them. What about the gateways? Forbes ran a piece this month on agent gateways, one governed hop that routes every model call, filters every tool, meters every token, a whole tier selling just the meter. And Palantir, whose paper argues correctly that the control layer must be model agnostic and owned by the institution with Palantir's platform as recommended home. Forrester recommended and recognized the category in December and started scoring it. I believe the Palantir paper, and believing it means applying its own logic to the control layer itself. Model liquidity has a sibling, control liquidity. If you escape model lock-in by moving everything into a single proprietary control layer, you cannot leave. You've pushed the dependency up one level. Of these, I kind of trust Palantir the most because they're not in the model business and they're not selling you AI tokens. But let's combine some of these. I actually like a combination of Palantir's model and ServiceNow's tower. These two sound in concert closest to what I'm proposing, so let me separate them. A tower where every vendor's agents report into one platform is a hub, and what I'm proposing has no top. Hold every vendor on the list and anyone quoting anthropic benchmarks at you to the same standard. If we sign, can we leave? Every vendor selling you a control plane is a reason that you actually need a federated one. Your Microsoft agents will be governed by Agent 365. Your ServiceNow agents report to the control tower. Your CRM, your ERP, and your developer's coding tools all arrive with a control plane attached. The day you deploy your second agent from your second vendor, you have a federated control plane. Nobody actually chooses federation. You inherit it, and the only choice left is whether you design or you fall into it accidentally. I propose design. An accidental federation is five vendor planes, each governing its own island with no layered metering costs across all of them and no single ledger. A design federation uses the same components, assigns the domains deliberately, and puts economics and observability above all of them. One disclosure since this episode is about incentives, Paragon does not sell models and doesn't sell anyone's control plane as of this episode. We have no Position in which any AI lab wins. That's the seat this argument has to be made from because the argument is that nobody with a position should hold the whole thing. In fact, as Paragon, we always say to our clients, your data, your infrastructure, our expertise. So let's talk a little bit about who watches the watchers. There's a harder problem underneath all of this vendor sprawl, and the vendors actually admit it. AWS's own well-architected guidance for AgenTech AI says, direct quote, a control plane that fails takes every agent with it, end quote. And rates the risk high. AWS is warning you about the architecture that AWS sells. That's honest. It's also an attack surface. In March, security researcher Johann Reyberger built something he called Agent Commander, a working malicious command and control dashboard operating agents from three different vendors built through prompt injection, control plane weaponized. Oi vei. Paradoxically, a universal control tower where every agent from every vendor reporting into one platform is the biggest watcher anyone has ever proposed. The academic version says any control that lives inside an agent's own process is defeatable. Real control requires separation of concerns. An independent layer that mediates actions before they execute from the outside. So who will watch the watchers? Whatever the answer is, it cannot be a bigger watcher. The evidence built you eight domains. The failure modes of centralization will tell you how to rearrange them. The market already federated your control plane. Our approach makes the federation more deliberate. Eight independent control planes, each with its own authority over every agent action assigned on purpose instead of being inherited coincidentally. So let's review the eight. Identity. Who is acting? Human, agent, or service, and with what scope? Policy. Is the action allowed under business rules and regulation? Security. Is it safe? Watching for injection, exfiltration, or poison tools. Trust. How much confidence does the output deserve? And when does a human review? Routing. Which model or tool handles the task? Economics. What does it cost? And does the budget allow it? Data governance. What may the agent see and change? Observability. What happened and why? One ledger for finance, audit, and engineering. The federation part is the arrangement. Every agent action gets evaluated across the domains before it executes. Some veto, some approve, some route, some advise. Some are actually redundant. We are still going to be experimenting with this architecture before we before we gel it in finality, but this is kind of how we're thinking through it right now. The important thing is no single domain and no single vendor holds all of it. So there's no single point of failure. You can always upgrade and add things to it and swap things out. And nothing for the quote unquote agent commander, uh the the evil, the evil agent commander to capture. Second and third opinions are built in structurally. Your auditors know this as separation of duties. Security people know it as defense in depth. And good architects see this as separation or concerns or suck. The vendor control planes don't get thrown out, they become participants. Agent 365 can hold identity for your Microsoft side agents. Your control tower can watch what it's connected to for the service agents, while an independent layer holds economics and observability across the entire state. And it goes past the platform giants because the market is full of specialized control planes, each one excellent at the job and that they do. And we're going to see more of these startups on very specialized control planes that tackle some aspect of security or engineering or cost. We're going to see gateways that meter costs. We're going to see observability tools that tie spend to outcomes. We're going to see governance platforms that enforce policy, eval suites that score trust. You staff each domain with the best tool for that job or several, and the domains can theoretically check each other. Best of breed used to lose the suite because integration was the expensive part. Agents talk to everything now, which puts the best of breed back on the table, and the decomposition has witnesses. Singapore derived most of these domains from the government side, Palantir from the vendor side. An OWASP, Mitra, and CISA fill in the rest. Each maps to domains. None fits inside one vendor's product. Then run reversibility test across the design. Routing makes the model choice reversible per task. Trust, borrowing Palantir's branching idea, makes agent actions reversible. You can fork, validate, roll back so agents get wider surface area at lower risk. Federation makes the control plane vendors themselves reversible because no participant holds enough of the estate to make leaving unthinkable. The truly irreversible decisions, your data leaving a region, your know-how entering someone's weights, get routed to the slow path with human review, exactly where the Singapore framework puts its checkpoints. Fast lanes for reversible decisions and the slow path for irreversible ones, and a layer that knows the difference. One more reason the Federation earns its keep, the next two years of your AI program are experiments. You're going to be trying lots of different things. Sovereign and open weight models against the big labs. This quarter's model version against last quarter's, one harness design against another. Every experiment needs the same environment. Route the same task both ways, score both inputs, meter both costs, log both runs. That environment is the federation. Routing runs the trial, trust scores it, economics prices it, observability keeps the record. Without the federation, every experiment is a one-off project without a control. With it, experimentation is much easier. So where do you start economics and observability? Those two stop the bleeding, make everything else measurable, and they're the two domains no vendor plane will ever hold honestly across the whole estate because every vendor meters its own island. Meter first, then govern. If you're an operating partner, this lands harder for you than anyone else in the audience. Talking to our PE audience here, Grant Thornton surveyed PE firms on AI governance. 9% could demonstrate defensible governance within 90 days, the lowest of any sector, and 7% have a tested AI incident response playbook. Meanwhile, FTI shows an alpha tier forming, where real AI capability is driving 18% more AI-related exits. The spread between governed and ungoverned portfolios shows up at the exit. A diligence exercise you can run next week, count a portco's accidental control planes. Every vendor agent platform they've turned on is one. Then ask who meters across all of them. If nobody can answer, that's your result and finding. And the same Gartner release from the cold open that we had earlier estimated about 70% of agents on the market are agent washing, workflow automation wearing a costume. And the only way to tell the difference is the auto trail, which is the observability domain doing diligence. The liability also concentrates upward. Under the EU AI Act, penalties reach 35 million euros or 7% of global turnover, and Covington's analysis as a sponsor with decisive influence may force enforcement directly, with turnover potentially including the sponsor. AON treats AI claims claims as a live DNO exposure, and the SAC has settled AI washing cases. And the model should feel familiar because the PE firm already runs a federation, independent companies, central oversight, and checks and balances. The federated control plane maps to it exactly, with the firm holding policy, economics, and observability, standards portfolio wide, while each portco runs identity, routing, and data governance locally against its own stack. The governance model you need is the one you already operate. Three questions to take to your team or your portcos at the next ops review. One, can anyone tell me what single AI task costs us per workflow per outcome? The answer is we see the monthly invoice, your economics domain is in trouble. Two, if an agent went into a loop at two in the morning, what stops it? And when could we find out? The answer is the invoice, that's $3,700 in terms of that question, and it's asking for the economics and observability together. Three, how many control planes do we already have and did we design the federation or inherit it? Count the vendor agent platforms you've turned on. Each arrived with a control plane attached. If nobody came and they can't name who holds the authority across all of them, you have your answer and you've got your first project. In our last episode on AI sovereignty, I ended with a sentence that I'll repeat. Sovereignty is your ability to choose. Today completes it. The federated control plane is what makes the choice reversible, and reversible choices are the only ones you can afford to make at the speed that this AI market moves. You're already living in a federation. The question is whether you designed it. Now here comes the sentence from my notes. The control plane is the token cost controller, is AI sovereignty, is AI security, and is the ROI unlock. It's the same thing. Whoever in your organization currently owns that sentence owns whether your AI investment converts to Evita. If nobody owns it yet, that's the first control gap to close. I hope that you're outside on a walk right now. I hope you're enjoying it. The AI space will have changed by the time you get back. Thanks for listening today and cheers.