Cloud & Infrastructure • 8 hours ago • Shruti Das

For years, cloud infrastructure has been optimized around a relatively predictable assumption: applications may be distributed, dynamic and highly scalable, but their workloads are still fundamentally controlled by software engineers. An application receives a request, calls known services, accesses defined resources and eventually returns a result. Even when workloads are complex, their execution boundaries, dependencies and permissions can generally be understood in advance.
AI agents are challenging that model.
An agent does not simply process an input and return an output. It can reason through a task, decide which tools to use, retrieve information from multiple systems, write or execute code, call APIs and continue working until it believes an objective has been achieved. As enterprises move from experimenting with individual agents to deploying fleets of them, infrastructure teams are discovering that the traditional cloud stack was not designed around this kind of workload.
That is why agentic infrastructure is rapidly becoming a cloud infrastructure problem rather than simply an AI application problem.
Recent developments across Kubernetes and the major cloud platforms point in the same direction. Google Cloud has been explicitly preparing its infrastructure and orchestration stack for what it describes as a surge of “agentic fleets,” with emphasis on scalability, isolation, security and cost management. Kubernetes, meanwhile, is developing Agent Sandbox specifically because conventional Kubernetes primitives do not map neatly onto long-lived, stateful and autonomous agent workloads.
Agents Are a Different Kind of Cloud Workload
The first mistake enterprises can make is treating an AI agent as just another microservice. A conventional microservice normally has a relatively well-defined contract. Platform engineers know which APIs it needs, what data it should access, how much CPU and memory it generally consumes and which other services it is allowed to communicate with. Those assumptions make familiar infrastructure mechanisms such as containers, Kubernetes namespaces, RBAC, network policies and service identities remarkably effective.
An autonomous agent behaves differently because some of its behaviour is determined at runtime. The agent may decide that completing a task requires access to a database, an internal API, a browser, a code execution environment or another agent. It can also change its execution path based on what it discovers along the way. This makes the infrastructure challenge less about keeping a service available and more about controlling what an autonomous workload is allowed to do while it is running.
That distinction is important because it changes the role of the platform team. The question is no longer simply, “Can we run this workload reliably?” It becomes, “Can we give this workload enough freedom to be useful without giving it enough freedom to become dangerous or unnecessarily expensive?”
Kubernetes Has an Abstraction Gap
Kubernetes remains an obvious foundation for agentic infrastructure because enterprises already use it for scheduling, networking, storage, identity and workload management. However, simply placing an agent inside a conventional Deployment does not solve the underlying problem.
Kubernetes’ own Agent Sandbox work illustrates why. The project describes agents as workloads that can be stateful, long-lived, mostly idle between tasks and capable of executing potentially untrusted code. That creates requirements around persistent identity, isolation, lifecycle management and rapid suspension and resumption that are awkward to express using traditional Kubernetes primitives alone.
This is where a new infrastructure abstraction begins to emerge. Instead of thinking about an agent as a pod that happens to contain an LLM application, platform teams may increasingly treat the agent execution environment itself as a first-class infrastructure object. That environment could include compute, persistent state, identity, networking, tool access and security boundaries, all managed as a single unit.
The implication for platform engineering is significant. The next generation of internal developer platforms may need to provide an “agent golden path” alongside existing application golden paths, with pre-approved runtimes, policies, observability and infrastructure templates designed specifically for autonomous workloads.
The Security Boundary Is Moving
Containers transformed cloud infrastructure partly because they provided a convenient unit of packaging and isolation. For many traditional applications, container isolation is sufficient when combined with operating-system and Kubernetes security controls.
AI agents complicate that assumption because they may dynamically generate and execute code. An agent could install a package, interact with a shell, access a tool or attempt a network connection that its developer did not explicitly anticipate. That means the runtime environment needs to be treated as part of the security architecture rather than simply as a place where application code happens to execute.
This is one reason microVMs and stronger sandboxing mechanisms are receiving renewed attention. The emerging model is not necessarily “containers versus virtual machines,” but rather containers for application packaging and stronger isolation boundaries for autonomous execution. Recent cloud-native discussions around agent sandboxing are already pointing toward combinations involving technologies such as gVisor, Kata Containers and Firecracker.
For infrastructure architects, this creates a new design question: which workloads are trusted enough to share a conventional container boundary, and which autonomous workloads deserve a stronger runtime boundary? That distinction could become as important as deciding between serverless, containers and virtual machines was in earlier generations of cloud architecture.
Identity Will Become Infrastructure, Not Just Security
The identity problem becomes even more complicated once agents start acting on behalf of people and systems.
A traditional service account might have permissions to read a database or invoke an API. Giving those same persistent credentials to an autonomous agent, however, can create an enormous blast radius if the agent is manipulated, compromised or simply makes an incorrect decision. The emerging architectural direction is therefore toward agent-specific identities and narrowly scoped, short-lived permissions.
Instead of asking whether an agent is “trusted,” infrastructure should increasingly determine what that agent can do right now, for this particular task, against this particular resource. Tool access becomes a policy decision, not merely an application configuration.
That also means identity systems, API gateways and policy engines will become increasingly intertwined with agent runtimes. The infrastructure layer needs to understand not only who initiated an action, but which agent performed it, which tool it invoked, what authorization it received and what happened afterward.
This is already reflected in enterprise platforms. Google Cloud’s current agent platform architecture, for example, combines agent identity, centralized tool access, policy enforcement and observability as part of its runtime environment. Microsoft’s Foundry Control Plane similarly focuses on tracing agent runs, controlling tool calls and applying runtime policies across agent fleets.
Observability Has to Follow the Agent’s Decisions
Traditional infrastructure observability tells engineers whether a system is healthy. Metrics show CPU and memory consumption, logs reveal application events, traces expose service dependencies and alerts identify failures. For agents, that is necessary but insufficient. An agent can be technically healthy while producing a disastrous result.
It may have successfully called every API, consumed an acceptable amount of compute and generated no infrastructure errors while still choosing the wrong tool, accessing inappropriate data or pursuing an incorrect course of action. Infrastructure teams therefore need visibility into agent behaviour as well as infrastructure behaviour.
This introduces a new layer of operational telemetry: tool calls, execution paths, task outcomes, policy interventions, token consumption, agent-to-agent interactions and potentially the reasoning context that can safely be captured.
The economics matter here too. An autonomous agent that loops for several hours may not trigger a conventional infrastructure failure, but it can quietly generate significant compute and model costs. Agent observability therefore becomes closely connected to FinOps.
In other words, the future cloud operations dashboard may need to answer questions such as: Which agents are running? What are they doing? What tools are they invoking? How much are they costing? Which policies are stopping them? And which workflows are repeatedly failing?
That is considerably more sophisticated than monitoring pods.
The Control Plane Becomes the Strategic Layer
As agent adoption increases, enterprises are likely to discover another problem: agent sprawl.
Different teams will build agents using different frameworks, models and cloud services. Some will run on Kubernetes, others on managed cloud runtimes, while others may exist inside SaaS platforms. Without a common management layer, the organization can quickly end up with dozens or hundreds of autonomous identities, inconsistent policies and fragmented observability. This is where the concept of an enterprise agent control plane becomes particularly important.
The control plane does not necessarily need to host every agent. Instead, it can provide a consistent layer for registration, identity, policy, tool access, observability, evaluation and governance across different execution environments. That is a familiar pattern in infrastructure. Kubernetes itself separates application workloads from the control plane responsible for scheduling and managing them. Cloud management platforms similarly abstract infrastructure operations away from individual servers.
Agentic systems may require the same separation: agents perform the work; the control plane determines the boundaries within which they are allowed to work.
This is becoming a significant enterprise architecture discussion, with industry research also highlighting centralized identity, policy enforcement, visibility and governance as necessary as organizations move from isolated pilots toward large-scale agent deployment.
What Platform Engineering Teams Should Do Now
The organizations that handle this transition well will probably not begin by rebuilding their entire infrastructure around AI agents. A more practical approach is to treat agents as a new workload category and gradually extend the existing platform engineering model.
That means establishing a few foundational capabilities before agent deployments become widespread:
- Dedicated agent runtime patterns: Define approved environments for running autonomous workloads instead of allowing every development team to invent its own architecture.
- Stronger workload isolation: Evaluate whether conventional containers provide an appropriate security boundary for agents that execute generated code.
- Agent identity: Give every agent a distinct identity rather than hiding autonomous activity behind generic service accounts.
- Fine-grained tool permissions: Limit which APIs, databases and external systems each agent can access and under what conditions.
- Agent-aware observability: Capture execution paths, tool calls, failures, resource consumption and policy interventions alongside conventional infrastructure telemetry.
- Cost controls: Establish budgets, quotas and automatic suspension mechanisms for agents capable of running for long periods.
- Human escalation paths: Define which actions require approval and which can happen autonomously.
- Standardized golden paths: Make the secure approach the easiest approach for developers to adopt.
The objective should not be to slow down agent adoption. It should be to prevent every business unit from independently solving the same infrastructure, security and governance problems.
The Cloud’s Next Abstraction May Be the Agent
Cloud infrastructure has repeatedly evolved by creating abstractions that hide underlying complexity. Virtual machines abstracted physical servers. Containers abstracted operating-system environments. Kubernetes abstracted cluster-level workload orchestration. Platform engineering then began abstracting those capabilities further through internal developer platforms.
AI agents may trigger the next major shift. The important question is unlikely to be whether enterprises will run agents. Increasingly, the question is what infrastructure abstraction will allow them to run thousands of agents safely, efficiently and predictably.
That is why agentic AI deserves attention from cloud and infrastructure leaders even when they are not building AI models themselves. The model may provide the intelligence, but infrastructure determines where that intelligence can act, what it can access, how long it can operate and how the organization can ultimately control it.
The cloud was built to scale applications. The next phase will require it to scale autonomous actors.
Key Takeaways
- AI agents are becoming a new class of cloud workload, with different requirements from conventional microservices.
- Kubernetes remains foundational, but traditional primitives are not enough for long-lived, stateful and autonomous agents.
- Runtime isolation is becoming a central infrastructure concern, particularly when agents can generate and execute code.
- Agent identity and fine-grained permissions will become critical, because persistent credentials can create an unnecessarily large blast radius.
- Observability must move beyond infrastructure health to include agent actions, tool calls, execution paths, cost and policy interventions.
- Enterprise agent control planes are emerging as a way to centralize governance across heterogeneous agents and infrastructure environments.
- Platform engineering teams have an opportunity to establish secure agent golden paths now, before agent sprawl becomes another infrastructure management problem.
