Skip to content
Some content is members-only. Sign in to access.

Bull vs. Bear: Can NVIDIA Hold Its Edge as Hyperscalers Go Custom?

Near-term NVIDIA demand stays strong, but custom accelerators give AWS and peers the leverage to cap pricing power by 2030.

By KAPUALabs

The contest over AI infrastructure is entering a more consequential phase. Hyperscale cloud providers are no longer merely purchasing accelerators from merchant vendors; they are designing their own. Amazon Web Services is the most advanced and vocal participant, but Google, Microsoft, and Meta are pursuing similar strategies. Alongside AMD’s expansion across CPUs and inference-optimized accelerators, this proliferation of proprietary silicon threatens to broaden the field of alternatives to NVIDIA and, over a three- to five-year horizon, erode its pricing power and market share.

The central question is not whether custom chips will immediately displace NVIDIA GPUs. They will not. The more important question is whether hyperscalers can use proprietary accelerators to reduce their dependence on NVIDIA, improve inference economics, and acquire greater command of the value chain. In industrial terms, the customers are becoming partial competitors. That is a structural development, even where the first deployments remain captive to the owners’ clouds.

AWS: Custom Silicon Moves from Experiment to Strategic Pillar

Amazon’s commitment is unequivocal. Jeff Bezos has designated custom chips as a new business pillar alongside AWS, Prime, and Marketplace 8, while the Graviton and Trainium product lines have developed into a substantial internal and commercial technology business 8. This is not a laboratory exercise conducted at the margin of AWS’s infrastructure strategy. It is a deliberate effort to control an increasingly important productive asset: the compute underlying cloud AI services.

Trainium3 reached general availability in late 2025, with UltraServers scaling to 144 chips, each carrying 144 GB of HBM3E 10. Nearly one million Trainium2 processors are already training and serving Anthropic’s Claude models 26. Adoption of Amazon’s custom silicon extends beyond headline AI deployments: 98% of AWS’s top 1,000 EC2 customers use Graviton 19. Multi-year, multi-gigawatt commitments from Anthropic and OpenAI provide further validation at frontier-laboratory scale 19. Reported across multiple sources in late July and early August 2026, these figures establish that AWS’s accelerators have achieved material scale and external credibility.

The commercial logic is straightforward. By controlling chip design, cloud deployment, and customer distribution, AWS can optimize the system rather than merely purchase one component of it. That combination may improve cost, performance, capacity planning, and bargaining power. It also creates long-term compute partnerships: once a customer has committed to substantial capacity and adapted its workloads to a platform, switching becomes expensive and operationally disruptive.

Complement Today, Merchant Threat Tomorrow

The immediate threat to NVIDIA is more measured than the headline numbers suggest. Amazon has publicly downplayed direct substitution, presenting Trainium as a complement to NVIDIA GPUs within a multi-platform portfolio rather than as a replacement 25. AWS continues to offer a broad range of NVIDIA instances, including B200, H200, A100, V100, and other configurations, and benefits from reselling NVIDIA-based compute 14,18.

The captive nature of proprietary accelerators also limits their immediate effect on the merchant market. Customers generally cannot rent TPU or Trainium as bare, interchangeable silicon; they must access these processors through the owner’s cloud 11. This arrangement gives hyperscalers a powerful internal differentiation tool without yet creating a fully open alternative to NVIDIA’s merchant platform.

The decisive issue is whether that boundary holds. Amazon is discussing selling Trainium outside AWS data centers 19. If realized, the move would transform Trainium from an internal cost and differentiation advantage into a direct merchant competitor to NVIDIA and AMD 19. The impact would not appear overnight. Multiple sources place the principal competitive effect over a three- to five-year period 19. Nevertheless, even chips that remain internal can weaken NVIDIA’s negotiating leverage by giving AWS a credible alternative in capacity planning and procurement.

The Hyperscaler Arms Race

AWS is not alone. Google’s TPU v7, known as Ironwood, Microsoft’s Maia 200 and Maia 300, and Meta’s MTIA 300 are all active or forthcoming custom accelerator programs 1,2,3,5,10,11,13,17. Major AI laboratories are also increasingly pursuing proprietary silicon to manage token-generation costs 15. Custom accelerators are therefore being treated not merely as engineering projects but as potential innovation moats 17.

This is the familiar pattern of industrial integration. In earlier eras, control of raw materials, transportation, and production determined who captured the surplus. In AI, the relevant combination is chips, compilers, models, data-center capacity, and cloud distribution. A hyperscaler that controls these layers can tune the entire system to its own workloads and customer base. The master resource is no longer simply access to a GPU; it is the ability to produce and allocate computation at a competitive cost.

Yet the resulting market will not be uniform. Inference is likely to be more fragmented than training because models, token-generation requirements, and response-time constraints favor different architectures 7,26. General-purpose GPUs will remain valuable where flexibility, software compatibility, and rapid deployment outweigh the benefits of specialization. ASICs will gain ground where workloads are stable and volume is sufficiently high to justify dedicated design. The likely outcome is coexistence, but coexistence under increasingly contested pricing and bargaining conditions.

AMD: The Broadening Alternative

AMD adds a second source of pressure on NVIDIA. It supplies CPUs to AWS 20 and has expanding relationships with Microsoft Azure, Google Cloud, Meta, and Anthropic 21. At the same time, it is acquiring Taalas to integrate model-specific inference silicon into its Instinct and Helios roadmaps 23,24. Its collaboration with Cerebras on ultra-low-latency inference using wafer-scale engines and the Helios platform further diversifies the field 9,16,22.

AMD’s strategic position is consequently unusual: it is both partner and rival to the hyperscalers. Its ability to offer a combination of conventional GPUs and specialized inference silicon gives customers a broader alternative to NVIDIA while also addressing workloads that might otherwise migrate directly to custom ASICs. This hybrid model may prove important as buyers seek to balance flexibility against unit economics rather than commit every workload to a single architecture.

Implications for NVIDIA

For NVIDIA, hyperscaler custom silicon represents a classic challenge from well-funded customers that are becoming partial competitors. The short-term impact remains limited by two protections. First, proprietary chips are largely captive to their owners’ clouds. Second, NVIDIA’s CUDA ecosystem retains considerable developer mindshare and creates meaningful switching costs. These advantages give the incumbent time, but not immunity.

The more serious risk is a ceiling on future pricing power. NVIDIA’s data-center business benefits from strong demand, scarce capacity, and the strategic importance of its accelerators. As hyperscalers deploy internal alternatives, they can negotiate from a stronger position even when they continue buying NVIDIA GPUs. Amazon’s AI revenue grew 37% year over year, driven in part by NVIDIA-based services 4. Amazon does not sell AI tokens; it earns infrastructure revenue regardless of the underlying chip 6. Thus, Trainium adoption does not automatically translate into lost AWS revenue or an immediate collapse in NVIDIA demand. It does, however, allow Amazon to retain more of the economics and reduce its dependence on a single supplier.

The multi-year, contracted nature of AWS’s AI capacity provides visibility 19, but it also signals that switching costs are high. Those commitments temporarily protect incumbent platforms while making the eventual architecture choices more consequential. If Trainium and comparable accelerators remain internal, they will still constrain NVIDIA’s leverage inside the largest clouds. If Amazon or another hyperscaler begins selling silicon beyond its own data centers, the threat will move from procurement pressure to direct merchant competition.

NVIDIA therefore cannot rely on a one-size-fits-all strategy. It must defend the CUDA platform while continuing to improve performance and economics across both training and inference. The strategic choice is between deeper vertical integration and a more open, customizable platform approach. Either path must answer the same industrial question: if a customer controls the cloud, the workload, and increasingly the accelerator, what prevents the merchant supplier from becoming a replaceable component?

Strategic Outlook

Amazon’s Trainium family has achieved meaningful internal scale and external validation. Anthropic and OpenAI serve as anchor customers, while the potential expansion into merchant sales could directly challenge NVIDIA’s accelerator business 19. The competitive effect is likely to develop over three to five years, with near-term disruption moderated by captive deployment models and continued complementary use alongside NVIDIA GPUs 11,19.

AMD’s acquisition of Taalas and its partnerships with Cerebras and the hyperscalers position it as a more comprehensive alternative for inference workloads, increasing pressure on NVIDIA’s data-center GPU segment 12,22,24. At the same time, NVIDIA retains substantial defensive strength through its software ecosystem, installed base, and ability to serve fragmented workloads. Those advantages provide a buffer, not a guarantee. The long-term erosion of pricing power and market share remains a material risk 7,19,26.

The durable conclusion is clear. Custom silicon will not end the merchant GPU market, but it will change its terms. The hyperscalers are building their own railroads into the AI economy, and NVIDIA will no longer control every route to compute. Its future position will depend on whether its platform moat remains strong enough to justify the premium—or whether specialized, integrated alternatives can move far enough down the cost curve to make that premium untenable.

Comments ()

characters

Sign in to leave a comment.

Loading comments...

No comments yet. Be the first to share your thoughts!

More from KAPUALabs

See all
| Free

Risk Factors Assessment

By KAPUALabs
/
| Free

Technical and Market Structure Analysis

By KAPUALabs
/
| Free

Regulatory and Legal Environment

By KAPUALabs
/
| Free

Market Sentiment and Analyst Coverage

By KAPUALabs
/