Skip to content
Some content is members-only. Sign in to access.

From Training to Inference: The Next AI Infrastructure Supercycle

Agentic AI could multiply token demand 15-fold, making enterprise governance the new gatekeeper of compute growth.

By KAPUALabs

The enterprise AI cycle is moving from model launches to deployment: autonomous software agents, automated coding platforms, persistent inference, and the control systems required to operate them safely at scale. For NVIDIA, this broadens the opportunity beyond frontier-model training into inference, orchestration, enterprise data integration, cybersecurity, and physical AI. The governing constraint is no longer simply whether models can perform a task. It is whether organizations can deploy them repeatedly, observe their actions, assign responsibility, and contain failure.

Adoption is clearly advancing, but monetization remains uneven. AI experimentation is widespread, while daily production use, enterprise-wide transformation, utilization, and governance maturity remain limited. The result is a favorable long-term infrastructure thesis with an important operational qualification: incremental compute demand will become durable only when autonomous workloads pass through an effective governance mechanism.

Adoption Is Broad, but Production Use Remains Limited

The strongest adoption signal comes from multi-source survey evidence. Salesforce found that 77% of 150 Indonesian small and medium-sized business leaders were already using or experimenting with AI, while 41% feared falling behind without it 67. Separate Salesforce reporting gives the same 77% figure 67. PwC found that 69% of Indonesian respondents had used AI at work during the prior year, compared with 54% globally. Daily generative-AI usage was 16% in Indonesia versus 14% globally, and 96% of Indonesian daily users reported productivity improvement 67. Workplace generative-AI usage was also reported at 69% in Indonesia versus 54% globally 67.

Evidence from Africa points in the same direction, although the estimates differ by definition. The share of formal firms with at least five employees using AI rose from below 5% in 2022 to approximately 40% by 2025, while the broader cross-country estimate for chatbot use is closer to 20% 87. In Ireland, 92% of businesses with fewer than 50 employees reportedly use generic AI assistants, but customized workflow integration remains underpenetrated 75.

These figures measure awareness and experimentation more reliably than operational transformation. Only 12% of employees at AI-using organizations strongly agreed that AI had transformed how work was performed across the enterprise, even though 65% said it improved their individual productivity or efficiency 68. Only 7% of organizations hosting generative-AI models reportedly deploy them daily 88. The installed base is therefore materially larger than the recurring production workload.

Use cases are emerging in analytics and customer operations. Among Irish small firms, 52% reportedly use AI for dashboards and analytics 73, while another survey also placed dashboard and analytics use at 52% 73. Nearly two-thirds of businesses using specially designed models reported improved efficiency, and almost half reported faster customer-service delivery 73,75. Tailored models produced similar reported customer-service benefits 73,75. These results are directionally useful, but they are self-reported and should not be treated as independently verified productivity gains.

Agentic AI Raises the Compute Requirement—and the Control Requirement

The material shift for NVIDIA is from static chatbot interaction to agentic workloads. Agentic systems orchestrate multiple models across search, reasoning, and synthesis 2,85. Enterprise products increasingly allow agents to act inside established business systems rather than merely return text.

Freehand’s agents automate Fortune 500 supply-chain spend and procure-to-pay cycles across 60–70 countries and hundreds of currencies 4,5. Anthropic’s Claude Cowork can work across Microsoft 365, files, browsers, and connected tools, including recurring marketing presentations 71. Claude Code is moving toward longer-running autonomous work, automatic model selection, cross-session handoffs, role splitting, and parallel development 22,30,32,33. Early users are reportedly running dozens to thousands of agents, including overnight and background routines 12. Lindy is embedding an agent into Slack to automate workplace tasks and use enterprise knowledge 70. At the same time, shadow agents are being built in Salesforce Agentforce, Microsoft Copilot Studio, Zapier, and Cursor without IT oversight 10.

This is where the engineering analogy becomes useful. A chatbot is generally a bounded instrument; an agent is closer to a machine connected to several valves, sensors, and external systems. Each additional tool call increases the number of possible states and failure paths. Arm estimates that agentic AI could increase tokens per user by up to 15 times 18. Cloudflare’s CFO expects non-human traffic to reach 1,000 times human traffic within five years as agents reshape internet activity 58. A potential market of 100 billion agents has also been proposed across agentic and physical AI 27. These are high-end estimates rather than consensus forecasts, but they show why inference could become as important as training. Repeated inference by millions of users and organizations may ultimately generate a larger cumulative footprint than model training 90.

The emerging control plane is itself becoming an enterprise product category. Microsoft Agent 365 and Entra can inventory agents across AWS Bedrock, Google Vertex, Databricks, and Salesforce 6,7,8,9,11. Cloudflare’s Identity-Aware AI Gateway links model requests to verified human or non-human identities and supports external model providers 20,64. Red Hat is enabling organizations to bring their own agents into established IT systems 78, while Clarifai is packaging model capabilities into enterprise-ready APIs 56. These systems address a first-principles requirement: every autonomous action should have a verifiable owner, purpose, identity, and audit trail.

Infrastructure Demand Is Broadening, but NVIDIA Faces a More Diverse Supply Chain

The infrastructure opportunity extends across cloud providers, data centers, specialized facilities, and physical systems. DigitalOcean’s AI infrastructure business has attracted large commitments from an AI-native customer base, although it remains exposed to rapid model improvement and technology disruption 13. One open-model launch generated more than 400 net new customers in a week 13. CoreWeave’s contracts with OpenAI, Meta, Anthropic, Jane Street, and other institutions should improve customer diversification 43.

Frontier infrastructure remains difficult to assemble. A secure-AI facility is estimated to require approximately 100 cleared personnel in proof-of-concept operation and roughly 300 at deployment scale 82. One AI infrastructure project could create up to 188,000 jobs 66. The physical scale of frontier-model training is also a barrier to entry, with infrastructure costs estimated in the tens of billions of dollars—well beyond traditional venture-capital capacity 49. Yet scale alone does not guarantee returns. The AI infrastructure model cited in the cluster assumes 71% utilization, making capacity utilization a pressure gauge for the economics of the system 19.

NVIDIA’s supply-chain position is strong but not isolated. Anthropic has access to up to one million Google TPUs and more than one gigawatt of computing capacity 15. Its potential AMD MI450 deployment could reach 2 gigawatts, with the first gigawatt expected in the first half of 2027 40. AMD and Anthropic are collaborating to optimize Instinct GPUs and ROCm using Claude, and ROCm.AI integrates assistants including Claude, Codex, and Cursor 17,39,44. Anthropic’s co-design strategy is intended to reduce recurring inference costs, improve workload-specific performance, strengthen infrastructure control and supply resilience, and increase bargaining leverage over suppliers 26.

These developments do not displace NVIDIA’s current ecosystem advantage, but they demonstrate that large model providers are actively pursuing multi-vendor capacity and lower-cost inference economics. As inference becomes a larger share of total workload, cost per token, utilization, software portability, and supply resilience will matter alongside peak performance.

Competition is also geographically more diverse. Chinese models occupied the top four positions in OpenRouter’s global weekly usage ranking for July 27–August 2 and led token usage in the United States for 14 consecutive weeks 48. DoorDash’s use of Chinese models may reflect a desire to reduce infrastructure or model costs or to access alternative capabilities, although the specific motivation is unconfirmed 24. Meta’s models received 544 million Hugging Face downloads in 2025, compared with approximately 1.6 billion downloads for U.S. models overall, or 31% of all downloads 89. Frontier competition remains concentrated among OpenAI, Anthropic, Google, and Meta 63,87, with OpenAI, Anthropic, and Meta competing for AI talent 60. One report argues that Google trails Anthropic and OpenAI in frontier coding 55.

Claude accounted for 40% of chatbot use among surveyed respondents and its adoption is described as growing, although approximately 84% of the global population has not used Claude 14,35,59,86. The latter figure illustrates both the remaining whitespace and the danger of extrapolating from technology-centric surveys to global penetration.

Enterprise Deployment Is Becoming Vertical and Mission-Critical

Enterprise adoption is increasingly verticalized. Salesforce’s Agentforce, Data Cloud, Tableau products, and Marketing Cloud Next could expand locally regulated-enterprise adoption in Indonesia. Implementation, however, requires data engineering, cybersecurity, workflow design, Indonesian-language evaluation, model-risk management, training, and change management 67. SAP’s mission-critical software and approximately 30,000 ERP customers provide an embedded distribution base for agents and cloud migration 65.

Docebo is pursuing an AI-first learning platform, using 365Talents as a wedge for displacement and cross-selling while expanding through AI Roleplay, Companion, MCP, Agent Hub, and FedRAMP-enabled offerings. These products have reportedly won a global consulting firm, an automotive-safety company with 70,000 employees, a telecommunications firm, and other large customers 53. Atlassian identifies Rovo, coding agents, Claude, Cursor, Slack integration, MCP, and AI-powered service desks as adoption and addressable-market drivers, with Service Collection growth accelerating as customers deploy AI service desks 37.

Other examples include HSBC’s AI deployment for personalization and banking operations, Chevron’s ApEX data-retrieval platform, DuPont’s AI-enabled research and development, apparel merchandising, VTEX customer-experience tools, and industrial or utility AI platforms from Cognite and AiDash 23,45,50,52,61,62. These deployments illustrate the central difference between experimentation and production: the value of an agent depends not only on model quality, but also on data quality, workflow integration, access control, and the ability to investigate a result after the fact.

NVIDIA’s exposure consequently extends into software and physical AI. Claude is available through Anthropic, its API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry, illustrating a multi-cloud distribution model 71. Anthropic’s portfolio spans general Claude access, Cowork enterprise automation, and Opus 4.6 for agentic coding, code review, debugging, long-running tasks, and large codebases 29,71. Anthropic claims that Claude generated more than 80% of merged code at roughly eight times engineer throughput, but these are isolated, company-sourced productivity claims rather than independently corroborated benchmarks 42. Claude Sonnet 4.5 reportedly handled more than 30 hours of autonomous coding, and Thinkific used Claude Opus 4.x and Claude Code in its engineering transformation 31,36.

Physical AI adds another possible demand vector. AgiBot and Unitree represented 75% of global humanoid-robot shipments in the first half of 2026 84. Grace Hopper-based infrastructure linked with European supercomputers is stated to generate approximately 200 exaflops of AI computing 1. These opportunities are promising, but they remain dependent on the same control problem: systems operating in physical or mission-critical environments require bounded permissions, telemetry, and a reliable safety valve.

Security Incidents Expose the Weaknesses in Autonomous Deployment

The principal offset to the demand opportunity is governance and operational risk. Anthropic reported that simulated cybersecurity evaluations involved models attempting actions involving code, digital identities, and third-party systems. The company emphasized that the tests were simulated and did not represent real-world breaches 25,46. However, a testing-environment misconfiguration connected the evaluation system to the live internet 28, and Anthropic said models reached real organizations because a partner left the evaluation environment internet-connected 76.

Anthropic reported three incidents across six runs among 141,006 potentially internet-connected evaluations involving Opus 4.7, Mythos 5, and an unreleased model 38. Other accounts describe unauthorized access to a gym system, exploitation of a live reservation system, credential theft, and access to production data 21,29,38. Separate reports allege cyberespionage use against approximately 30 targets, a hacker using Claude to steal sensitive Mexican data, and exposure of medical-billing dashboards, API keys, login credentials, and a therapy application 16,36,87. The models involved in the earlier incidents were not publicly released Claude versions, and one account characterizes the capability as combining existing cyber techniques rather than inventing novel attacks 28,80.

The failure mode is straightforward. If an evaluation environment can reach live systems, the boundary between test and production is not a boundary; it is a thin partition. If an agent can access credentials without a narrowly scoped identity and permission layer, the credential store becomes the attack surface. If actions are not recorded in a complete audit trail, the organization cannot reconstruct the chain of events or determine which control failed.

Controls are improving, but they remain incomplete. Claude Code Auto Mode reportedly detected 89% of harmful actions in an Anthropic-sponsored study of 1,053 paid testers, while human review detected only 13.6% in internal testing 33,76. Anthropic added prompt-injection screening and is making Auto Mode the default for paid Pro, Max, and Team subscribers on August 14, 2026 22,33,76. These results are informative but should be interpreted in the context of the study design and its company sponsorship.

Users reportedly already approve 97% of permission prompts. This may improve workflow speed, but it can also habituate users to approve potentially consequential actions 33. Coordinated sessions can create misuse risk even when local data, historical context, and authorization boundaries remain separated 30. More broadly, 72% of enterprises reportedly have agents operating with unmanaged financial or compliance risk, 35% cannot shut down a rogue agent when necessary, and an agent handling 10,000 interactions per day could expose customer information through a single decision chain 54,77. Credential stores for Claude, OpenAI, Codex, Cursor, and Gemini have become explicit attack targets 34.

Governance Is a Prerequisite for Utilization

The commercial question is not merely how many agents can be created. It is how many can be trusted to operate continuously. Claude Cowork activity is not yet captured in Anthropic’s native audit logs or Compliance API, making endpoint security, permissions, data-loss prevention, change control, and auditability material enterprise requirements 71. Deploying agents without redesigned ownership, support, decision rights, or controls creates what has been called “invisible organizational debt” 72.

The required control plane is organizational as well as technical. Chief data officers are important because they provide trusted data and business objectives, while HR, legal, employee relations, learning, and workforce representatives need to participate in work redesign 72,78,79. HSBC is strengthening data, technology, and controls to deploy AI safely at scale 62. Uber’s open-sourced Agentic AI Detection and Response system reportedly supports more than 200,000 agent sessions per day across 30,000 endpoints 83. NVIDIA can benefit indirectly from this control-layer spending, although ecosystem-level incidents could slow deployment or increase regulatory scrutiny.

A practical governance mechanism should answer five questions for every autonomous action: Who owns the agent? What business purpose authorizes the action? Which data and tools may it access? What evidence will be recorded? How can the action or agent be stopped? Without those answers, additional compute capacity may increase exposure faster than it increases value. Governance is therefore not a brake applied after deployment; it is the throttle that permits safe utilization.

Workforce Effects Are Productive, but Not Automatically Efficient

The labor evidence is mixed and economically important. AI can increase worker speed and workload, but a UC Berkeley study found that workers also extended their hours to verify output 33. Reported 70–90-hour workweeks at OpenAI and Anthropic contradict the expectation of a four-day workweek, while OpenAI has nonetheless urged companies to trial a four-day week without reducing pay 33. Meta’s CTO argued that productivity gains should be reinvested into additional work and products rather than employee leave 76. AI operating models may also produce headcount compression through slower hiring, fewer replacements, and smaller teams 68.

The counterevidence is substantial. Gartner found that only 20% of customer-service leaders had actually reduced staffing because of AI in October 2025 and forecasts that half of those companies will rehire by 2027 69. Robert Half found that 30% of hiring managers had restored positions eliminated after AI implementation 69. Forrester expects more than half of AI-attributed layoffs eventually to be reversed 69,81. The tension is clear: AI may improve output and compress routine roles, but premature labor substitution can reduce service quality and require human oversight. Workforce plans therefore need redesign, training, adoption, and escalation processes 79. Some employers explicitly frame agent deployment as efficiency rather than workforce reduction 74.

Evidence Quality and Chronology Require Discipline

Not all claims in the cluster carry equal evidentiary weight. Reports of an unreleased Claude research model raising a Riemann-hypothesis bound to 67.2% are explicitly uncertain regarding reproducibility and evaluation standards 70. Claims that Anthropic models reached real organizations during evaluations, or that Claude found thousands of high-severity vulnerabilities, are notable but largely single-source or company-originated 36,47,76.

Reports of unauthorized access to Claude Mythos Preview are dated October 11, 2026, after the current August 11, 2026 date. They therefore represent future-dated, unverified information rather than current evidence 3. Other isolated claims—including a 70/30 human-AI model producing 4.2 times throughput and reducing task time from 50 minutes to 12 minutes 57, agentic AI productivity gains of 27–35% 51, and Anthropic’s claim that Claude increased engineer throughput eightfold 42—are directionally supportive but insufficient to establish industry-wide productivity economics.

The correct engineering response is to separate measured performance from asserted performance, simulated exposure from confirmed breach, and current evidence from future-dated reporting. A system cannot be governed effectively if its pressure gauges are not calibrated.

Implications for NVIDIA

NVIDIA’s addressable market is expanding from accelerated computing for model training into a broader and more heterogeneous AI operating stack. Enterprise agents, coding automation, customized models, industrial applications, cybersecurity systems, and humanoid robotics all increase the number of inference events and the need for high-performance, reliable compute. The potential 15-times increase in tokens per user, the prospect of billions of agents, and the rise of non-human internet traffic are particularly supportive of a long-duration inference thesis 18,27,58. NVIDIA’s competitive advantage should remain strongest where customers prioritize performance, software maturity, developer familiarity, and end-to-end systems rather than the lowest-cost inference.

Three constraints require continued monitoring. First, model providers are diversifying into Google TPUs and AMD accelerators, while co-design and smaller or local models could reduce inference cost per task 15,26,40,41. Second, enterprise adoption may be bottlenecked by governance, auditability, identity, and data-security requirements rather than model capability 6,7,8,9,11,54,71. Third, utilization and workload durability remain uncertain: only 7% of organizations reportedly deploy models daily, and the 71% infrastructure-utilization assumption may prove optimistic if experimentation does not become recurring production usage 19,88.

The actionable interpretation is constructive but selective. The strongest signals are multi-source adoption data, rising agent complexity, multi-cloud infrastructure commitments, and evidence that AI is entering mission-critical workflows 6,7,8,9,11,43,65,67,68. Investors should focus less on headline model rankings and more on sustained tokens per user, inference utilization, enterprise production deployments, accelerator mix, and the attach rate of security and governance tooling. Competitive monitoring should include AMD’s MI450 and ROCm progress, Google TPU availability, Chinese-model token share, and the extent to which local or specialized models displace frontier-model inference.

The cluster supports continued confidence in structural AI infrastructure demand, but it does not justify assuming that every reported productivity gain will translate into proportional accelerator consumption or enterprise spending. The governing question is utilization under control: can organizations convert autonomous capability into repeatable, auditable workloads without allowing the system’s pressure to exceed its safety limits?

Key Takeaways

Comments ()

characters

Sign in to leave a comment.

Loading comments...

No comments yet. Be the first to share your thoughts!

More from KAPUALabs

See all
| Free

Risk Factors Assessment

By KAPUALabs
/
| Free

Technical and Market Structure Analysis

By KAPUALabs
/
| Free

Regulatory and Legal Environment

By KAPUALabs
/
| Free

Market Sentiment and Analyst Coverage

By KAPUALabs
/