Okta Targets AI Agent Token Costs With Identity-Based MCP Tool Scoping
Okta's new MCP tool-scoping feature cuts visible tools by up to 90%, shrinking the 'tool tax' that inflates every AI agent's token bill.
Every time an AI agent makes a model call, its prompt gets stuffed with a menu of every tool it might use — even the ones it will never touch. That menu is rendered from the schemas, names, descriptions, and parameters of every tool an MCP server exposes, on every single turn. And the model has to reason about all of them, whether or not it ever calls them.
Okta calls this the “tool tax.” On August 12, 2026, the identity security company published a detailed analysis showing how its identity-based MCP tool-scoping feature can cut the number of tools visible to a model by more than 90%, reducing the token costs tied to tool descriptions by roughly the same margin. The approach reframes AI agent cost optimization as an identity and authorization problem rather than a gateway or metering problem.
The tool tax explained
The proliferation of Model Context Protocol (MCP) servers has been one of the fastest-growing patterns in enterprise AI. MCP provides a standardized way for AI agents to discover and call external tools — databases, APIs, SaaS integrations, internal systems. But as organizations deploy more MCP servers and expose more tools, a hidden cost emerges.
Each tool exposed by an MCP server contributes a JSON schema (name, description, parameter definitions) to the model’s prompt on every turn. A single popular MCP server might expose dozens or even hundreds of tools. Multiply that across thousands of active users, each running agents that make multiple turns per task, and the token costs compound rapidly — before any actual work is done.
The problem has been largely invisible to AI teams because it does not surface as a rejected API call or a security incident. It appears as a token bill that is harder to explain than it should be. “More tools exposed per server means more tokens spent per agent decision, on tool selection alone, before any actual work happens,” Okta explained in its analysis.
How identity-based scoping works
Okta’s solution operates at the point where an agent connects to an MCP server. Within the Okta dashboard, an administrator configures which tools a given identity — whether a human user or an AI agent — is entitled to use. When the agent requests the tool list from the MCP server, Okta returns only the scoped subset, stripping away everything the agent is not authorized to access.
This shorter list gets injected into the agent’s prompt on every turn. The result: less token spend per turn, before the agent ever attempts a tool call. Unauthorized tools simply do not exist in the list the model sees. Scope is then checked again at runtime before any tool call executes, providing defense-in-depth.
The mechanism is an extension of Okta’s existing AI agent identity platform, “Okta for AI Agents,” which reached general availability in April 2026. That product replaces hardcoded credentials and standing access with scoped, short-lived tokens issued only for what an agent needs. The tool-scoping capability builds on this foundation by applying the same least-privilege principle to the tool catalog itself.
Results from internal modeling
Okta’s internal modeling, published alongside the announcement, tested the approach across a realistic mix of permission levels. The company defined several representative user segments — helpdesk read-only, helpdesk operator, app admin, brand and email admin, and super admin — and weighted each by an assumed share of monthly traffic.
The findings were striking. In some scenarios, scoping cut the number of visible tools by more than 90%. Because tool-schema token cost tracks tool count nearly linearly (each tool contributes roughly the same amount of schema text to the prompt), the cost reduction followed suit. A helpdesk read-only agent that previously saw 100 tools might now see only 5, with a proportional drop in per-turn token consumption.
Okta was careful to note that absolute savings depend on average schema size, request volume, and model pricing, all of which vary by deployment. The company reported results as percentages rather than dollar figures. However, for organizations running agents at scale across hundreds or thousands of users, even a 50% reduction in tool-schema tokens could translate to significant monthly savings.
Why this is an identity problem, not a gateway problem
A key argument in Okta’s analysis is that cost control for AI agents is fundamentally an identity and authorization challenge, not something API gateways can fully solve.
Gateways — whether purpose-built AI gateways or general-purpose API management tools — excel at metering spend by API key, team, or group. They can cap spending after the fact and provide visibility into consumption patterns. But they operate on group-level information, which cannot determine what a specific agent or the person behind it is actually entitled to access.
“Capping spend after the fact makes whoever owns the gateway the token police almost by default,” Okta noted. Someone has to set limits, defend them when teams push back, and handle escalations when users hit caps mid-task. Gateways meter what already happened; they cannot prevent decisions from being expensive in the first place.
The identity layer, by contrast, has per-user and per-agent entitlement data. Okta filters the tool list at a resolution that gateways simply do not have access to. The gateway still meters what gets through, but the identity layer shrinks what there is to meter — so there is less to police.
Paul Webber, Principal Cybersecurity Industry Analyst with Software Analyst Cyber Research, endorsed this approach: “Cost control for agents is best provided using identity governance tools that offer more granular control and precision without disrupting business processes. Okta’s approach is an elegant way to do this because it leverages the same entitlement data that governs security, not a separate metering layer without that insight.”
The security angle: smaller blast radius
The cost savings are only half the story. The same scoping that keeps the prompt lean also keeps the blast radius small. By stripping away tools that an identity is not authorized to use, the organization simultaneously reduces token costs and the potential damage an attacker could inflict if that identity were compromised.
This dual benefit aligns with Okta’s broader “blueprint for the secure agentic enterprise,” which asks organizations running AI agents to answer three questions: Where are my agents? What can they connect to? What can they do? The tool-scoping capability is a direct implementation of the second question, answered at a finer grain than most teams currently check.
Least-privilege MCP reduces both the tokens a model has to reason about and the tools an attacker could exploit. It is a rare case where the security-optimal choice is also the cost-optimal choice.
Broader context and implications
Okta’s announcement arrives at a moment of intense industry focus on AI agent economics. As agentic workloads move from demos to production, the cost of running autonomous agents — which make many more model calls than simple chat interfaces — has become a top concern for enterprise buyers. MCP has emerged as the de facto standard for connecting agents to external systems, with support from Anthropic, OpenAI, Google, and others.
The tool-tax problem is not unique to Okta’s platform. Any organization using MCP servers with large tool catalogs faces the same overhead. What Okta brings is a governance framework that ties tool visibility to existing identity and access management policies, avoiding the need for a separate metering or filtering layer.
As agent deployments scale, the gap between organizations that govern their tool catalogs proactively and those that do not will widen rapidly. Okta’s data suggests that the difference can be measured not just in security posture but in hard token costs — up to 90% of which may be spent reasoning about tools that agents never actually use.
Sources
- [1] https://www.okta.com/newsroom/articles/fewer-tools-fewer-tokens-how-okta-cuts-agent-tool-selection-costs-before-they-happen/
- [2] https://www.artificialintelligence-news.com/news/okta-targets-ai-agent-token-costs-with-mcp-scoping/
- [3] https://www.okta.com/blog/product-innovation/agent-gateway-runtime-governance/