Dharma Insights — Operational№ 027 · Infrastructure
← The Signal№ 027 · Infrastructure · July 21, 2025 · 5 min read

AI Agents Compute

AI Agents Are Breaking Compute: Why a New Infrastructure Stack Must Rise The landscape of artificial intelligence is rapidly evolving, with AI agents emerging as a transformative force. These autonomous…

AI Agents Are Breaking Compute: Why a New Infrastructure Stack Must Rise

The landscape of artificial intelligence is rapidly evolving, with AI agents emerging as a transformative force. These autonomous entities, capable of perceiving, reasoning, and acting , are pushing the boundaries of what computing infrastructure can handle. The consensus among industry experts is clear: traditional compute models—whether on-premise, cloud, containerized, or serverless—are fundamentally mismatched to the unique demands of AI agents, leading to economic inefficiencies, technical limitations, and operational inadequacies. A revolutionary rethinking of infrastructure, tooling, cloud models, and energy sources is not just an option, but a necessity to unlock AI’s potential.

The Core Mismatch: AI Agents vs. Legacy Compute

Existing compute paradigms were predominantly designed for predictable, stateless, CRUD (Create, Read, Update, Delete) applications. AI agents, however, operate with a distinct set of characteristics that expose the shortcomings of these legacy systems:

  • Variable Compute Intensity: AI agents experience heavy inference bursts, followed by periods of idle states..

  • Persistent State: Unlike stateless functions, agents require context to survive across interactions, demanding persistent memory.

  • Specialized Hardware Demands: GPUs and TPUs are crucial, but often needed in bursts rather than continuously allocated.

  • Bursty Scaling: Their usage is typically tied to human interaction or specific real-world events, necessitating dynamic scaling.

  • Latency Diversity: Real-time tasks (like chat) coexist with deferred tasks (like background analysis).

  • External API Dependencies: Agents frequently rely on multiple external service layers for real-time knowledge fetches and actions.

This fundamental mismatch results in significant inefficiencies, including the overprovisioning of compute resources and wasted idle GPU cycles. Traditional autoscaling tools, built for predictable flows, simply fail to map to AI's multi-phase workload patterns (context-fetch → inference → action). The verdict is clear: cost models designed for CRUD logic break down when applied to AI agent compute economics.

The Data Center Bottleneck: A Looming Crisis

The AI boom has triggered an unprecedented supercycle in data center investment. Yet, this growth is facing severe, inherent constraints:

  • Energy and Cooling: AI workloads are projected to consume a massive amount of global electricity (10% by 2030) , a demand that could surpass that of entire nations like Brazil. Cooling infrastructure and generator deliveries are already experiencing significant delays, ranging from 5 to 24 months.

  • Supply Chains and Design Obsolescence: Parts shortages are lengthening deployment timelines. Furthermore, AI accelerators like NVIDIA H100s are evolving at a pace that outstrips data center construction, meaning facilities planned today may be technically obsolete by the time they launch.

Without a move towards "reconfigurable, interoperable, energy-conscious" data center designs, scaling AI effectively remains a pipe dream. Data centers are becoming strategic weapons, but only if built with vertical integration, energy intelligence, and AI-native orchestration in mind.

Cloud Models Are Splintering: The Rise of AI-Native Platforms

Hyperscale clouds (AWS, GCP, Azure), while powerful, were not built with AI's real-time, context-rich workflows in mind. They often present high costs for GPU-heavy workloads, offer difficult infrastructure abstraction, and provide inflexible inference runtimes.

This environment has created fertile ground for new "AI-native" and developer-centric cloud builders. Platforms like Runpod and Together are emerging as leaders, offering GPU-native, developer-first UX tailored specifically for AI agents. The future of cloud is envisioned as:

  • Vertically Integrated: Abstracting away infrastructure complexities, akin to platforms like Netlify/Vercel for GenAI.

  • Developer-Centric: Offering easy inference orchestration and transparent costs.

  • Edge-First: Pushing compute closer to users for real-time loops and lower latency.

This shift suggests that hyperscalers may lose ground to more agile, AI-native cloud providers, or be forced to reinvent their offerings.

Salesforce's AI Vision: The Frontend of the Agent Stack

Salesforce, a leader in enterprise solutions, is rapidly advancing the "top of the AI stack". Their initiatives, as highlighted in the "Great India Sales, Marketing & Service Summit", showcase a commitment to AI-driven transformation:

  • Salesforce Data Cloud: Unifies structured and unstructured data for "real-time, hyper-personalized engagement" by creating a comprehensive unified customer profile.

  • Einstein GPT + Prompt Engineering: Democratize AI across marketing, sales, and service functions. Prompt engineering is recognized as a crucial lever for "Unlocking Workflow AI Potential" and optimizing AI outputs to achieve "Precision inputs [to] optimize Einstein GPT outputs".

  • MuleSoft & Slack GPT: Orchestrate multi-agent workflows, decisions, and actions across disparate systems , forming a "foundation for AI Agents" through API-led connectivity and intelligent automation.

  • Tableau Pulse: Turns metrics into real-time insights and actions, empowering data-driven decisions at the flow of work.

However, this sophisticated AI vision inherently demands a backend compute infrastructure that is equally flexible, memory-aware, and agent-optimized. Salesforce's pursuit of context-rich, real-time, personalized AI experiences highlights that the underlying compute must support bursty GPU inference, state retention, and personalized context orchestration at scale. The current reality is that while the cognitive interface for enterprise AI (e.g., Salesforce's offerings) is rapidly maturing, the "cognitive infrastructure" to back it is brittle and lagging.

The Path Forward: A Compute Reformation

The AI wave is not merely a software revolution; it is a profound compute reformation. To fully unlock the value of platforms like Salesforce and to deliver real-time, personalized, intelligent AI at scale, we must reinvent compute from the bottom up. This necessitates a multi-faceted approach:

  • Agent Architecture: Designing agents to support bursty compute, persistent state, and elastic GPU access.

  • Infrastructure Strategy: Moving towards elastic, GPU-aware platforms with fine-grained autoscaling. Building agent-first platforms with observability for reasoning and memory.

  • Data Centers: Shifting toward energy-optimized, reconfigurable, and AI-native designs. This includes co-developing infrastructure with energy providers (e.g., nuclear, hydrogen) and leveraging modular solutions.

  • Cloud Models: Transitioning from general-purpose clouds to vertically integrated, developer-first AI platforms.

  • Enterprise Stack Alignment: Ensuring that Salesforce-like intelligent workflows are seamlessly aligned with agent-native backend infrastructure.

The winning stack for the AI agent era will be memory-rich yet compute-light when idle, GPU-accessible but not GPU-bound, edge-ready, and cloud-fluid. It will be architected not around traditional functions, but around feedback loops and reasoning cycles. Those who continue to treat infrastructure as a mere "container for models" will inevitably fall behind. The time to rethink compute is now.

View all signals →