The relevant subject is broader than NVIDIA alone. It is the evolving AI-infrastructure ecosystem in which Alphabet competes, supplies demand, and may face disintermediation. The central pattern is clear: demand for AI compute remains structurally strong, but economic advantage is moving beyond the sale of accelerators toward control of the complete stack—custom silicon, memory, advanced packaging, networking, power, cooling, software, cloud capacity, and application-specific inference.
For Alphabet, this creates both scale advantages and strategic vulnerabilities. Proprietary chips can reduce dependence on NVIDIA, yet rapid hardware cycles, supply constraints, capital intensity, and falling compute prices can weaken the returns on infrastructure investment. We must therefore distinguish between the short-run equilibrium, in which leading-edge capacity is scarce and NVIDIA commands substantial pricing power, and the long-run adjustment, in which hyperscalers develop custom silicon, suppliers expand capacity, and alternative architectures become more credible.
The strongest conclusions are supported by multiple sources. Jensen Huang’s position as NVIDIA’s chief executive is confirmed by seven sources 5,6,7,8,9,29,45, while NVIDIA’s CUDA advantage over AMD is supported by three sources 51. NVIDIA’s H100 rental economics, Samsung’s HBM5 and base-die plans, and Supermicro’s ten-GPU systems also have meaningful multi-source support 1,3,26,60,78. By contrast, many product-performance claims, rumored transactions, financing concerns, and geopolitical allegations are single-source or explicitly unconfirmed. They are best treated as scenario inputs rather than established facts. The source material is concentrated in July 2026, with the latest claims extending to August 1, 2026; it therefore describes a rapidly adjusting market rather than a settled long-run equilibrium.
Key Insights
The bottleneck is broadening beyond GPUs
The evidence supports a continuing shortage of AI infrastructure, but the shortage is no longer confined to accelerators. NVIDIA’s chief executive argues that the semiconductor industry remains far too small and faces shortages in memory, storage, optical interconnects, packaging, and foundries 79. Semianalysis and DS Investment & Securities reportedly see shortages continuing beyond 2030 25. NVIDIA’s order books are described as full 48, AI leaders are requesting materially more memory than previously forecast 30, and NVIDIA is seeking additional memory from suppliers 30.
The supply response is necessarily gradual. Spending cannot remove GPU scarcity within a single fiscal year 27. HBM suppliers require years to add capacity, reach acceptable yields, and qualify products 56, while new entrants face billions of dollars of investment, specialized supply chains, testing requirements, and customer-qualification barriers 56. Samsung believes a meaningful increase in industry supply before 2028 is unlikely 68, although memory capacity could eventually overshoot demand 70.
For Alphabet, the immediate implication is that infrastructure availability—not merely model quality—can constrain AI-product deployment and cloud monetization. GPU-dense facilities may require hundreds of megawatts to gigawatts of power 38, and hyperscalers must now assess power and cooling before deploying each new generation of silicon 87. The broader cloud and GPU market is constrained by electricity generation, transmission, grid connections, semiconductor manufacturing, and geographic concentration 72. Liquid cooling is consequently becoming an important opportunity at the intersection of high-performance computing and AI 10, while the semiconductor build-out benefits industrial-gas suppliers such as Linde 53. The Genesis Mission could add demand across accelerators, networking, memory, fabrication equipment, cloud and supercomputing, scientific software, laboratory automation, and electricity services 57.
The supply chain is becoming more geographically distributed, but it remains dependent on a small number of critical nodes. TSMC has outlined a $100 billion U.S. expansion 24 and a broader $265 billion investment plan 20. Production in Arizona is expected to progress from 4nm to 2nm by 2028 20, although management has indicated that the schedule could remain flexible and extend over as much as ten years depending on market conditions 24. TSMC says it is responding to long-term structural growth rather than surrendering opportunities to competitors 20. Its chief executive has dismissed bubble concerns while urging disciplined investment 33. TSMC’s N2 mass production was scheduled to begin in 2025 4,13, although helium shortages could delay 2nm production 33. TSMC reportedly has approximately five times Intel’s advanced-packaging capacity 13, and Intel could not immediately absorb all of NVIDIA’s packaging demand 13. The strategic constraint is therefore packaging as much as wafer fabrication.
NVIDIA’s asset-light model provides flexibility and supports margins, but it leaves the company dependent on external production, packaging, assembly, and geopolitical continuity 69. It relies heavily on TSMC and other manufacturers 69, with reported concentration in China-based assembly 69 and exposure to Taiwan, China, and international supply chains 17,69. The specific claim that NVIDIA secured 50–60% of TSMC’s advanced CoWoS capacity and HBM supply years in advance is single-source 22. It is nevertheless directionally consistent with the broader, multi-source evidence of packaging and memory scarcity.
Custom silicon is the central competitive question
Hyperscalers are developing custom chips to reduce reliance on NVIDIA 2,56, and numerous hyperscalers are reportedly building accelerators around HBM 56. This is the central strategic tension for Alphabet. Internal silicon can lower cost per token, improve workload-specific efficiency, and protect capacity availability. It also transfers capital requirements, engineering work, supply commitments, and execution risk to the hyperscaler. Custom silicon is therefore eroding NVIDIA’s moat without yet eliminating it 75.
NVIDIA still controls more than 95% of data-center GPUs according to the Futurum Group, compared with approximately 4.5% for AMD 51. CUDA remains substantially more advanced than AMD’s software ecosystem 51. The competitive threat is nevertheless becoming more credible at the system level. AMD’s Helios is claimed to offer 15% greater AI compute and 50% more memory than NVIDIA’s Vera Rubin NVL72 41, while AMD claims up to 30% more tokens per dollar and 50% more memory 41. AMD’s stated goal is to take share from NVIDIA 42, but its ability to displace NVIDIA at the system level remains uncertain 40. A modeled 20–25% GPU share is consequently a scenario rather than an established outcome 51.
Supermicro’s support for dense 5U systems containing up to ten AMD Instinct MI350P GPUs 78 demonstrates increasing platform choice, although product availability alone does not establish durable ecosystem share. The relevant elasticity of substitution is not uniform: a customer may replace a GPU for a particular workload more easily than it can replace the surrounding software, networking, memory, support, and system-integration capabilities.
NVIDIA is responding by broadening beyond GPUs. It sells GPU chips 32, supports any model architecture that generates additional chip demand 32, and benefits from both closed and open models 79. Huang argues that open models increase usage, which in turn drives demand for computing hardware and data-center capacity 46,79, while NVIDIA believes the market has misread the effect of free models 79. NVIDIA is developing the 550-billion-parameter Nemotron Ultra model 14, has added PhysicsNeMo and CUDA-X to its software ecosystem 82, and is expanding into agentic computing through Vera 73. Vera also challenges Intel and AMD in data-center CPUs 73, consistent with NVIDIA’s renewed CPU focus in 2026 51 and its earlier server-CPU offering 51. NVIDIA has developed a custom CPU 73, while Vera is positioned as an expansion into agentic computing 73.
Alphabet should therefore be viewed not only as a purchaser of NVIDIA hardware but also as a silicon and platform competitor. The movement from search toward generated answers is a structural change affecting NVIDIA 79 and Alphabet’s core business model. Alphabet’s opportunity is to convert AI infrastructure into differentiated products, search economics, cloud services, and enterprise workflows. Its risk is that infrastructure becomes increasingly commoditized while model and application competition capture the returns.
Memory and advanced packaging are strategic constraints
Memory content is rising rapidly across AI systems. The H100 uses five HBM stacks 56. Rubin, B200, and GB300 systems are described as using eight stacks, with B200 carrying 192GB of HBM 56. NVIDIA’s next-generation Feynman chip is expected to use at least 16 stacks 56, and successive NVIDIA generations are using increasing numbers of HBM stacks 56. Bullish investors argue that memory content rises exponentially with each generation 55. A GB300 NVL72 server is reported to contain more than 20 terabytes of HBM 34.
This pattern supports a favorable medium-term outlook for memory suppliers. Micron is described as the fastest-growing HBM supplier, producing HBM3E and volume-producing HBM4 for NVIDIA’s Rubin platform 56. It is building additional plants and increasing spending 15. Micron and SK Hynix produce premium HBM for NVIDIA 28, while Samsung is pursuing a 2nm GAA base die for HBM5 26. Samsung’s prior Tesla wins and broader orders from Tesla, AI, HPC, and cloud-service-provider customers demonstrate process and customer traction 26. Samsung is reportedly negotiating five-year supply agreements with global data-center operators 62, and Micron’s commentary supports a thesis of longer contracts and less-cyclical business conditions 55.
The bullish memory thesis has an important counterforce. NVIDIA may reduce HBM per rack and pursue memory pooling or system-level approaches through CMX to control costs 56. AMD’s claimed memory advantage could increase competitive pressure 41. If NVIDIA or hyperscaler orders weaken, all three major HBM suppliers could be affected 56. Supplier leverage is therefore high, but it is not immune to a slowdown in cloud capital expenditure. Samsung also faces execution risk in adopting 2nm by 2028 26. Chinese memory producers remain materially behind in EUV and may require a decade or more to approach Western competitors 61,71, while Chinese RAM is reportedly unlikely to reach global markets in meaningful volume for another three or four years 63.
For Alphabet, the issue is less direct HBM revenue exposure than the cost and availability of internal AI infrastructure. Greater memory content can improve model throughput and support larger workloads, but it raises bill-of-materials costs, power requirements, and dependence on concentrated suppliers. The strategic value of custom silicon will depend not only on its design, but also on whether Alphabet can secure memory and packaging on competitive terms.
Economics are shifting from scarcity pricing toward utilization
The evidence presents an instructive apparent contradiction. H100 rental prices increased approximately 40% between October 2025 and March 2026 1,3,60, and renting an H100-equivalent GPU is cited at $250,000 annually 47. Yet H100 availability has improved materially from two years earlier, when capacity was almost impossible to rent 19. Used H100 prices have fallen sharply 19, creating a purchase opportunity for lower-cost infrastructure 19. Customers can either buy hardware or consume capacity as an operating expense, while cloud providers monetize GPUs through hourly rental revenue 19.
The most reasonable interpretation is a bifurcated market. Current-generation capacity can command a premium even as older equipment depreciates. A100 GPUs remained useful and rentable approximately six years after launch 60. Used A100 80GB units sold for roughly $10,000, compared with an original retail price near $20,000 before ancillary components 60. H100 pricing has held up around four years after launch 60, but newer Blackwell systems have higher operating costs 60, and Rubin systems require still more power and cooling 60. Blackwell deployment also faces more construction delays 60.
The market is consequently moving from a singular focus on buying raw accelerator capacity toward extracting more useful work from existing hardware 49. LinkedIn reportedly doubled GPU efficiency and plans to keep GPU investment, compute, and storage capacity flat in fiscal 2027 47. A slowdown in model demand would be particularly damaging to cloud operators: utilization could fall while debt, leases, power, and other obligations remain fixed, reducing future NVIDIA orders as well 83. New generations may offer better performance per watt and lower cost per token before older systems earn their required returns 83.
This is directly relevant to Alphabet’s capital allocation. A model-driven increase in token demand supports cloud and AI investment, but efficiency gains can reduce the number of GPUs required for a given workload. Conversely, lower cost can stimulate usage sufficiently to raise total compute demand. Huang’s view that the semiconductor market must become five to ten times larger over ten years 79 is therefore a bullish demand thesis, not a guarantee of equivalent unit growth or pricing power. Claims that GPUs deteriorate in three years 31 and that accelerator chips become obsolete within a couple of years 60 represent conservative depreciation scenarios. The six-year usefulness of A100s and sustained H100 pricing provide counterevidence, while the expected operating life of some current-generation GPUs through 2030–31 59 illustrates the uncertainty.
Alternative architectures remain options, not established substitutes
The competitive field extends beyond AMD and hyperscaler ASICs. China is developing photonic computing through Guangdong’s 2024–30 action plan, which targets core technology, photonic computing, and on-chip optical neural networks 81, alongside an industrial cluster exceeding RMB100 billion 81. The optical-computing sector is expected to evolve as a hybrid ecosystem rather than a winner-take-all replacement 81. Commercialization still depends on improving precision beyond approximately eight bits, expanding matrix dimensions beyond 128×128, and unifying software 81. The forecast that photonic chips could reach 30% of intelligent-computing-center share within five years 81 is an ambitious single-source prediction rather than consensus.
Other alternatives include Cerebras’ wafer-scale engines for inference and throughput 58, a claimed “Frozen v2” efficiency improvement of up to ten times 18, and a Korean NPU claimed to deliver fourfold efficiency, although that claim is disputed 64. A hypothetical Korean design delivering 80% of NVIDIA’s performance at half the cost or power 64 is useful for scenario analysis but is not evidence of commercial displacement. Decentralized GPU networks could expand compute capacity if demand continues to exceed traditional supply 12, and DePIN demand could be catalyzed by decentralized GPU computing 84. Insufficient user demand remains a stated risk 77, while the proposed stack—distributed inference, proof of computation, token incentives, validator consensus, and monetization of unused GPUs—remains largely conceptual 77.
For Alphabet, these technologies represent option value and a reminder that centralized, CUDA-based infrastructure is not the only possible architecture. Low-cost legacy CUDA compute could open markets in drug discovery, protein folding, climate simulation, and computational fluid dynamics 60, extending hardware utility even as leading-edge systems turn over. The necessary distinction is between technical possibility and deployable infrastructure. Software compatibility, precision, reliability, supply, and total cost of ownership will determine whether these alternatives affect Alphabet’s economics.
Geopolitics and customer concentration create structural risk
Export controls have effectively removed China revenue from NVIDIA’s current forecast 79, and NVIDIA’s current assumptions are described as including no China revenue 79. Beijing’s ban on H200 orders, followed by limited exceptions, shows how national-security and self-sufficiency objectives can override customer preference and efficiency 52. The Trump administration’s subsequent loosening of controls for approved H200 sales triggered criticism from U.S. lawmakers 52. Tencent’s access to NVIDIA’s latest GPUs is constrained by geopolitical issues 35, while Chinese companies are shifting toward domestic alternatives 74.
Chinese chips may remain a generation behind, but the gap is narrowing 74. Domestic deployment gives Chinese companies guaranteed demand, operating experience, scale, and software-optimization opportunities 80. Chinese chip stocks rallied on expectations for Huawei accelerators 74, and China may be making progress in lithography and manufacturing 56.
The source material also includes more speculative allegations of indirect access to NVIDIA chips by Alibaba through a Southeast Asian third party 36. Alibaba denies the allegations, which conflict with NVIDIA’s stated position 36. NVIDIA is also the subject of a Taiwan AI-chip-smuggling investigation 50. These claims should not enter a base case without independent verification, but they illustrate the compliance complexity surrounding advanced AI hardware. Export restrictions can eliminate geographic revenue opportunities 79, and NVIDIA identifies export controls as a principal value risk 11. Similar restrictions could affect Alphabet indirectly by limiting hardware available to international cloud customers or by accelerating regional and sovereign-AI deployments.
Customer concentration compounds the risk. NVIDIA’s three direct buyers reportedly represent 54% of revenue, with additional indirect concentration 69. Apple reportedly redirected $50 billion that it declined to spend on GPUs toward NVIDIA 23, while NVIDIA has vendor-financing arrangements 85. The claim that NVIDIA may be financing or backstopping its own demand, with potential repossessions, a collapse in compute prices, and hardware write-downs if customers fail, is a bearish scenario drawn from a single discussion 65. It is not corroborated, but it identifies a genuine cycle risk: if infrastructure providers overbuild against speculative demand, falling utilization could impair cloud customers and chip suppliers simultaneously.
Ecosystem strategy strengthens the moat while increasing governance exposure
Huang is central to NVIDIA’s strategy and public narrative 79. He advocates broad consultation, open competition, rapid innovation, and application-led technology rather than fear-based restrictions 79. He has argued that regulation can be shaped by incumbents seeking competitive advantage 79. NVIDIA and 24 other companies supported targeted legal frameworks against unlawful extraction from closed models rather than broad restrictions on open-weight models 11. The company also argues that independent inspection and community improvement can strengthen trust and transparency 16.
NVIDIA reinforces its ecosystem through investments and partnerships. It has invested in or partnered with CoreWeave, which receives prioritized GPU shipments and support 21; Lambda, which gains access to newer GPUs 35; and Nebius, whose NVIDIA Cloud Partner status provides access to current generations 35. Nebius is scaling at triple-digit growth 90, although NVIDIA does not count the investment as revenue 14. A proposed NVIDIA–NAVER infrastructure project faces GPU and memory constraints, obsolescence risk, and regulatory or export-control changes 39. NVIDIA’s investment philosophy is to support a broad set of potential winners rather than select a single one, with ecosystem investments small relative to its resources 14.
The strategy carries financial and regulatory caveats. NVIDIA’s market position and customer relationships could attract regulatory scrutiny 69. Privacy, contractual data use, model safety, intellectual property, and regulation are material risks 79. Applications in medicine, transportation, autonomous vehicles, cybersecurity, and other regulated sectors create compliance and liability exposure 79. NVIDIA’s private, less-liquid investments may be difficult to revalue or exit during a downturn 86. The Groq transaction reportedly involved $13 billion due at closing and another $4 billion within one year 69, using licensing and mass hiring that may reduce regulatory scrutiny and disclosure 69. These are NVIDIA-specific risks, but they matter to Alphabet because the AI ecosystem increasingly combines hardware ownership, strategic investment, cloud contracts, and platform governance.
Demand is spreading into automotive, robotics, storage, and enterprise infrastructure
Demand for GPU-enabled systems is expanding across businesses, universities, governments, and cloud providers 88. Robotics requires accelerated GPU compute 43, while Toyota planned to adopt NVIDIA Drive AGX Orin and Thor 66. Thor configurations are described as exceeding 2,000 TOPS and supporting more than 20 sensors 66, although high-end automotive compute systems face potential supply bottlenecks 66. Semiconductor capacity is identified as a key factor in Toyota’s autonomous-driving prospects 66.
Storage and interconnects are also becoming strategic beneficiaries. Silicon Motion expects an NVIDIA ecosystem ramp and five additional Tier 1 customer ramps in the second half of 2026 91, with significant PCIe Gen 6 growth projected for 2028 91 and an initial addressable-market target of 15–20% 91. Its customer diversification spans GPUs, DPUs, TPUs, telecommunications, servers, flash manufacturers, and cloud providers 91. It nevertheless faces slower PCIe 5 client-SSD adoption and the risk that multiple vendors capture AI storage 91. The need for faster and higher-volume data flows rises as computing power expands 89.
Nutanix and AMD may benefit from rising token consumption, data sovereignty, enterprise-inference economics, private and hybrid-cloud AI, and demand for EPYC CPUs and Instinct accelerators 76. This is relevant to Alphabet’s Google Cloud strategy. Sovereign and private AI deployments may diversify demand away from centralized hyperscaler infrastructure, while enterprise inference may favor cost-efficient and distributed architectures. Nebius and Vultr illustrate continuing demand for high-end NVIDIA cloud capacity 35,37, and Nscale is preparing for a multibillion-dollar IPO 54. Specialized GPU clouds therefore remain investable even as hyperscalers develop internal silicon.
Implications for Alphabet
For Alphabet, the evidence points to a three-part strategic test.
First, Google must maintain access to leading-edge compute while reducing dependence on NVIDIA’s pricing, supply allocation, and software stack. The custom-silicon trend is favorable if Google can achieve better performance per dollar and per watt on its own workloads. It simultaneously threatens NVIDIA’s pricing power and may reduce the long-run value of generic GPU capacity.
Second, Alphabet must convert infrastructure scale into durable monetization. Rising model quality and open-model availability can stimulate usage and cloud demand 46, but efficiency improvements can flatten infrastructure spending, as illustrated by LinkedIn 47. The relevant measure is therefore not GPU deployment alone, but cost per token, utilization, and the incremental revenue generated by each additional unit of infrastructure.
Third, Alphabet must manage the physical and financial constraints of the data-center footprint. Power, cooling, memory, packaging, and construction schedules may become more binding than processor availability. The investment question is not whether AI demand is large, but whether demand grows quickly enough—and remains sufficiently profitable—to justify the capital committed before the next generation of hardware changes the equilibrium.
The broader infrastructure data reinforce the durability of the cycle but not the durability of excess returns for every participant. The market is transitioning toward system-level economics: tokens per dollar, performance per watt, utilization, memory bandwidth, networking, and software portability. NVIDIA retains a formidable position through CUDA, scale, product cadence, ecosystem partnerships, and system integration. Its moat is nevertheless exposed to hyperscaler ASICs, AMD system alternatives, photonics, inference-specific hardware, open models, and lower-cost architectures. Rapid innovation creates the possibility that a competing architecture or model reduces the value of existing NVIDIA hardware 79, while a more extreme tail scenario involves rapid substitution of CUDA or NVIDIA accelerators 65.
Alphabet’s relative position is strongest when it combines proprietary silicon, software, data, and distribution rather than competing solely for rented compute. The company benefits from the same secular expansion represented by the ambition for a five-to-tenfold increase in semiconductor-market size 79, the growth in server wafer content from 54 square centimeters to 110–130 square centimeters between 2015 and 2025 44, and the broader rise in silicon content across cloud, AI, automotive, and digital products 44. Global 300mm wafer demand grew at approximately 6% CAGR from 2000 to 2025 44, with servers representing 18% of 2025 end-use demand 44. The five largest wafer manufacturers serve approximately 75% of the market 44. These figures describe a durable infrastructure cycle, but not necessarily durable excess profits for all suppliers.
The principal downside case for Alphabet is not an abrupt disappearance of AI demand. It is a gradual reduction in the need for hyperscaler-scale infrastructure before recent investments earn acceptable returns. More efficient models, local models, custom silicon, and distributed compute could reduce hyperscaler data-center construction, GPU purchases, power generation, and AI-infrastructure investment 67. If model demand weakens, cloud utilization and chip orders could fall while fixed obligations remain 83. Conversely, if robotics, autonomous driving, scientific computing, cybersecurity, and enterprise inference expand more rapidly than efficiency improves, the infrastructure cycle could remain undersupplied for years.
Conclusion and Monitoring Priorities
The evidence supports a constructive but selective view of AI infrastructure. NVIDIA remains the benchmark and primary system-level competitor 40, supported by demand across open and closed models and by a deep software ecosystem. Claims of tenfold efficiency improvements, 30% better tokens per dollar, rapid photonic penetration, or imminent alternative-chip displacement are mostly single-source assertions and should not determine a base-case valuation.
For Alphabet, the actionable focus is narrower and more concrete: proprietary-accelerator adoption, cost per token, Google Cloud AI utilization, data-center power availability, HBM and packaging commitments, and evidence that AI products are generating incremental revenue faster than they are increasing capital intensity. The market will evolve through adjustment rather than a single discontinuity. Under current conditions, NVIDIA’s dominance remains substantial, but the long-run equilibrium will depend on how quickly substitution becomes practical, how efficiently hyperscalers deploy their own silicon, and whether AI demand expands faster than the cost of serving it.
Key takeaways
- AI-infrastructure demand remains structurally strong, but constraints now encompass memory, advanced packaging, power, cooling, networking, and grid capacity—not only GPUs 56,72,79.
- Hyperscaler custom silicon is the central strategic variable: it can improve Alphabet’s cost and supply position while weakening NVIDIA’s moat, although CUDA and system-level integration remain substantial advantages 2,51,56,75.
- Economics are shifting from accelerator scarcity toward utilization and efficiency. Falling used-GPU prices and improved performance per watt could pressure infrastructure returns even as aggregate AI usage rises 19,49,83.
- Alphabet should monitor proprietary-chip deployment, cost per token, AI-cloud utilization, power and memory procurement, and whether AI monetization is outpacing the capital required to support it.