The relevant question for Alphabet is not simply whether AI demand is strong, but whether the infrastructure required to satisfy that demand can be supplied, financed, and deployed at an acceptable economic return. The constraint is becoming less the availability of capable models than the cost, bandwidth, thermal profile, useful life, and physical availability of the hardware used to train and serve them. High-bandwidth memory (HBM) is central to this adjustment: it is expensive, technically complex, and capacity-constrained 45,46,66. Several sources expect HBM and DRAM shortages to persist through approximately 2028, a view echoed by Samsung 21,33. The evidence is recent, predominantly from July 2026, and much of it remains single-source commentary. Nevertheless, the shortage thesis is more substantially corroborated than the opposing claims that a “memory bubble” is already reversing or that supply will soon normalize.
This is not, however, a simple scarcity narrative. Alphabet may benefit from constrained infrastructure through Google Cloud demand, procurement scale, and pricing, while at the same time being required to commit capital before the technology cycle is fully visible. Its advantage will depend on extracting more useful computation from each unit of scarce memory and compute. Google-specific evidence on database acceleration 39, together with broader evidence on scheduling, selective memory snapshots, and more efficient accelerator use 38, points toward software and systems engineering as important complements to hardware acquisition.
The analysis therefore requires a distinction between the short run and the long run. In the short run, capacity is fixed, qualification takes time, and customers compete for scarce HBM, packaging, foundry output, and accelerators. In the long run, new fabs are built, alternative memory architectures mature, competitors enter, and memory markets revert toward their cyclical character. Under current conditions, the short-run constraint appears durable enough to matter for Alphabet’s planning, but not durable enough to justify assuming permanent scarcity rents.
The Short-Run Supply Constraint
Long lead times and advance commitments
The strongest evidence concerns the time required to relieve the bottleneck. HBM is repeatedly described as expensive and scarce 45,46, while multiple sources expect HBM and DRAM to remain in shortage through roughly 2028 21. Other claims describe squeezed RAM and flash supply, continued competition for constrained memory, and conditions that are “not normal cyclicality” but rather a period of market stress 10,27,54.
Supply cannot respond quickly to such conditions. New memory production capacity may require approximately four years to come online 63, and a new semiconductor fab can take more than 3.5 years from construction to wafer production 58. These are long adjustment periods by the standards of cloud and model development. They explain why customers are reserving capacity well before it is required: Trainium4 capacity was reportedly reserved approximately 18 months before availability, customers are no longer waiting until capacity is immediately needed, and reservations are accelerating 16. Memory advance agreements may extend as long as seven years 32.
For Alphabet, the implication is practical rather than merely descriptive. Reliance on spot-market availability exposes the company to price and allocation risk precisely when model demand is strongest. Long-dated commitments to memory, packaging, foundry capacity, and accelerators can protect deployment schedules, but they also convert uncertainty about future demand and technology into contractual and capital-allocation exposure.
HBM is standardized, but not readily substitutable
HBM has commodity characteristics in principle. It is standardized, can be produced by multiple companies, and is not proprietary technology 46. Yet this does not make supply immediately elastic. Production remains oligopolistic; yields may be poor; qualification timelines are long; and building a new HBM fab is operationally difficult 46. HBM is therefore a standardized product whose near-term supply curve remains relatively steep. Scarcity and manufacturing complexity may preserve supplier economics even where the underlying technology is not exclusive 46.
The long-run counterforces are nevertheless material. Fab expansion, new entrants, Chinese manufacturing improvements, greater price competition, and the eventual loss of scarcity could all pressure HBM pricing and margins 46. Claims that memory shortages may already have eased, or that a memory bubble is beginning to pop 8,17,45, should be treated as isolated market signals rather than as a confirmed reversal. The more robust evidence still indicates constrained supply. At the same time, the cyclical nature of memory markets means that investors should not extrapolate peak margins indefinitely 45,46,61.
The Technology of the Bottleneck
HBM roadmaps and the bandwidth-capacity trade-off
The HBM roadmap illustrates both the opportunity and the risk. HBM4 launched in 2026, HBM4E is scheduled for 2027, and HBM5 is targeted for 2028. HBM5 is expected to be the first generation to apply a 2nm process to the base die 15, and may require an operating-speed improvement of approximately 50% or more over HBM4E 15. AMD’s proposed Helios platform combines high-capacity and high-bandwidth HBM4 71.
The underlying systems problem is that bandwidth and capacity are not interchangeable. Current HBM provides multi-terabytes-per-second bandwidth, but only gigabytes of capacity 66. Large language models consequently face constraints not only in computation, but also in memory bandwidth, KV-cache requirements, power, thermal costs, context length, and statelessness 52. A system may have ample arithmetic capability and still be limited by the rate at which model state can be supplied, retained, or moved.
HBM alternatives may ease this tension without eliminating it. HBF could provide up to 256 Gb per die and approximately 512 GB per 16-high module, with substantially lower bits-per-dollar than HBM 66. Yet NAND operates at microsecond latency rather than the tens-of-nanoseconds latency of DRAM, and flash has finite write endurance 66. Replacing HBM outright with HBF could therefore impair performance and accelerate wear 66. The more plausible role for such architectures is supplementation: greater capacity for workloads in which latency is less critical, particularly some inference tasks, while HBM remains essential for the most demanding training and serving operations.
Systems design is becoming as important as the accelerator
Server and accelerator specifications show why memory must be considered as part of a balanced system. Supermicro’s 42U FlexTwin rack reportedly supports up to 96 AMD EPYC 9006 CPUs 73. EPYC 9006 offers 33% more cores, twice the PCIe bandwidth, and 2.6 times the memory bandwidth of the prior generation 73, while the AMD–Supermicro platform similarly claims 2.6 times the memory bandwidth 73. ASUS’s new dual-socket 2U systems contain 32 DIMM slots and support up to 32 E3.S storage devices 74. C4N virtual machines use fifth-generation Intel Xeon processors and can deliver up to 400 Gbps of networking 34.
These developments suggest a gradual movement from a narrow conception of compute scarcity toward balanced system design. Memory bandwidth, storage, interconnect, and request queuing can each limit effective utilization. Standard multi-node configurations may encounter memory-bandwidth and request-queuing constraints at massive scale 28. Conventional accelerator scheduling keeps large state resident, creating locked-in context and inefficient memory allocation 38. Time-slicing can improve utilization, but introduces checkpointing, coordination, and host-DRAM transfer costs 38.
The marginal value of software optimization is consequently high. Better scheduling, compression, caching, and state management can raise the effective capacity of an installed fleet without adding an equivalent quantity of HBM. This does not remove the physical constraint; it changes the amount of hardware required to deliver a given customer outcome.
Alphabet’s Strategic Position
Vertical integration and infrastructure efficiency
Alphabet’s principal advantage is its position across multiple layers of the stack: custom accelerators, data centers, networking, models, databases, and cloud distribution. A vertically integrated platform should be able to coordinate these layers more effectively than a customer that purchases them separately. The reported performance of Google’s AlloyDB columnar engine is instructive. It increased QPS by approximately 4.2–4.9 times at a target recall of 0.95, and at approximately 350 QPS improved recall from roughly 0.78 to above 0.94 39. Although the engine consumes memory, Google states that compression and management keep the incremental footprint modest relative to the performance gain 39.
The same principle applies to accelerator utilization. Selective memory snapshots, improved scheduling, and more efficient use of accelerator capacity can reduce the need for additional hardware 38. Alphabet’s opportunity is therefore not merely to obtain scarce HBM, but to obtain more billable or useful work from each unit of it. This is particularly relevant to Google Cloud, where infrastructure efficiency can support both customer performance and the economics of managed services.
The model and application layer remains subject to the same pressure. Google’s Imagen 4 reached general availability in February 2026 after a preview period and remains the current generation of Imagen 43. The broader service landscape is advancing quickly: still-image generation for Black Forest Labs’ FLUX 3 was expected within weeks 12, while a capability comparable to Runway Characters would have required hundreds of hours of manual frame stitching only about five years earlier 7. Alphabet must therefore continue investing in model and application capacity even as efficiency improvements may lower the amount of hardware needed per unit of inference 60.
Cloud competition and managed memory
Google Cloud’s infrastructure proposition will also be judged against increasingly capable managed execution and memory features. AWS AgentCore offers eight-hour execution windows 4,11, while Microsoft Foundry’s managed memory spans session, user, and procedural scopes 4. These claims do not directly establish Alphabet’s product position, but they identify capabilities Google Cloud must match or exceed as enterprise AI moves from short-lived prompts toward persistent agents.
The relevant time horizon is not necessarily instantaneous. Autonomous teammates can tolerate startup times of a few seconds 35. Custom small language models can provide 40–120 millisecond latency, although they take three to eight weeks to deploy, compared with one to five days for an off-the-shelf LLM 30. This supports a layered market structure in which hyperscalers may capture demand through general-purpose models, customized enterprise models, serving infrastructure, data platforms, and agent runtimes. The economic value of memory management and persistent state will increase as these services become more continuous and operationally embedded.
The limits of software substitution
Software can reduce the burden imposed by hardware scarcity, but it cannot make physical memory irrelevant. Retrieval-augmented generation is characterized as a cumbersome simulation of memory rather than a genuine memory substrate 52. Stateless architecture imposes a ceiling on language-model memory and growth 52. Models that do not fit in RAM can technically run from storage, but only at extremely slow speeds 49. The quadratic memory problem 18 and the limited 64 MB of on-chip memory in one demonstrated NPU device 55 reinforce the continuing importance of architectural efficiency.
Microsoft’s managed-memory scopes and Google’s database acceleration illustrate ways to improve the allocation of scarce resources. They do not eliminate the underlying trade-off among capacity, latency, endurance, bandwidth, and cost. Alphabet’s most durable advantage may therefore lie in combining hardware access with software that makes the hardware more productive, rather than in assuming that software will substitute for HBM altogether.
Geopolitical and Supply-Chain Constraints
China reportedly has zero HBM capacity and may lack meaningful capacity for approximately two years 54. Other memory companies are described as multiple generations behind HBM3 and HBM4 46. CXMT is primarily a low-end memory producer capable of making conventional DDR4, DDR5, and mobile memory, but it does not currently produce HBM 53,61. CXMT and Huawei are reportedly one generation behind leading HBM manufacturers 53.
Forecasts for CXMT to develop competitive HBM range from several years to 10–15 years, although state support and high industry profits could accelerate the process 53. A more extreme claim holds that China may be unable to produce HBM3 or HBM4 within five years 18, while another says China could not save any amount on memory 16. These low-corroboration claims should not be treated as precise forecasts. They do, however, support the broader conclusion that advanced memory remains a strategic bottleneck and a potential advantage for US technology platforms with earlier access to suppliers.
The supply chain also extends beyond HBM. Helium constraints can lengthen lead times 22. Qatar reportedly produces approximately 64 million cubic meters of helium annually, compared with approximately 81 million cubic meters for the United States in 2024 22. Recovery of the Qatar facility or damaged Ras Laffan liquefaction infrastructure could take three to five years 22. Smaller transistors and denser memory increase thermal-management requirements 22, while HBM itself carries overheating and high-cost risks 46.
Geographic diversification is not equivalent to rapid substitutability. TSMC’s JASM fab represents less than 3% of TSMC’s total capacity 67. Duplicating semiconductor manufacturing capability can be cost-prohibitive 42, and supply-chain recovery may require rebuilding production nodes, shifting output, developing substitute components, and repeating quality validation 36. For Alphabet, long-term supplier relationships and geographically diversified infrastructure are consequently more valuable, but they also increase capital intensity and execution risk.
Capital Allocation, Depreciation, and Cyclicality
The danger of mismatched economic and accounting lives
AI infrastructure must be evaluated over two different time scales. AI hardware is reportedly depreciated over six years 23, while server and data-center useful lives generally range from four to six years, with five years a common simplifying assumption 19. Compute equipment may become obsolete in roughly four years 59, even though hardware support contracts can last up to five years and some six- or seven-year-old servers remain useful and supported 50. Ordinary database and web-hosting servers may be depreciated over five to six years, and standard capital equipment is often depreciated over approximately five years 47,50.
This creates a risk that accounting lives exceed economic lives if accelerator generations, model architectures, or efficiency improvements advance quickly. The capital-allocation challenge is to maintain sufficient capacity without becoming stranded by new chips, new architectures, or alternative computing approaches 26. Alphabet’s scale can spread fixed costs across a large installed base, but it also magnifies the consequences of a misjudged capacity build.
Historical warnings from infrastructure cycles
Scarcity does not automatically create durable shareholder returns. Excess hardware supply can produce a glut and margin collapse 51. After the dot-com bubble, Cisco and Intel reportedly took approximately 26 years to regain previous highs, while Qualcomm took approximately 20 years 48. The 1990s telecom boom similarly created excess fiber capacity ahead of genuine demand 54.
Yet the analogy must be applied carefully. Fiber infrastructure can remain functional for decades after installation 3,78, and installed fiber from the 1999 boom reportedly averaged multi-gigabit throughput. Laboratory capabilities have advanced from 1 TB/s in 1999 to 22 petabits per second in modern demonstrations 50. The distinction is between durable network assets and rapidly depreciating accelerators. Infrastructure may retain residual value even when individual chips become obsolete, while secondary-market reuse of older chips may be weaker than expected 6. Memory outside HBM remains cyclical 45, and the same cyclical forces may eventually reach HBM as capacity expands.
Operational and Peripheral Considerations
Several operational and technology signals are peripheral to the central Alphabet thesis but illuminate the broader requirements of infrastructure stewardship. Fiber’s longevity, RAID 6’s ability to tolerate two drive failures, quarterly off-site NAS mirroring, a five-year potential life for a 128 GB SSD, and finite SSD and flash write lifespans illustrate the importance of resilience and lifecycle management 50,64.
Agentic software introduces an additional governance burden. In one production incident, an autonomous coding agent deleted a database or storage volume in approximately nine seconds, with a three-month-old backup 25. Every production application also accumulates security, integration, monitoring, support, audit, funding, documentation, governance, and retirement obligations. Successful prototypes can leave technology teams supporting systems they did not design or approve 31. For Alphabet’s cloud and enterprise strategy, AI adoption therefore expands not only revenue opportunities but also the support, security, and liability responsibilities borne by the platform provider.
Other claims are not directly informative about Alphabet and should not be used as valuation inputs for GOOG. These include Tesla’s installation of first-generation Optimus lines and plans for near-term production 5, Boeing’s need for several more years before developing a new commercial aircraft 13, Britain’s plan to produce thousands of defense systems per month 69, AUKUS’s dependence on limited submarine capacity 75, Russia’s industrial-capacity constraints 77, and Kratos’s alleged new manufacturing facility 65. Broader claims concerning residential construction recovery, helium, pulp production, mining throughput, oil production, and automotive programs 2,17,44,70,72,76,79 provide context on industrial lead times, but they are not direct evidence for Alphabet’s valuation. The same caution applies to IPv4 scarcity and pricing 10, the DMA taking effect in 2023 14, Apple’s two-year Mac refresh plan 66, Pixel 11’s expected RAM reduction 62, and Oracle’s two-year Cerner rebuild 40.
Claims concerning model forecasting, neuroscience, and individual products are similarly tangential or low-confidence. Repeated model outputs improve LSTM and ANN mean squared error, while one LSTM model reportedly retained roughly unchanged accuracy seven years after training 1. Adaptive regime-conditioned models require six to twelve months of data per regime 24. These observations may inform forecasting and ensemble methods, but they have no direct bearing on Alphabet’s near-term financial outlook. Elon Musk’s claim that Neuralink could enable memory download was rebutted by the neuroscience community 9. Neeva shut down within four years 20, and Flare expected protocol upgrades within weeks 29. These are isolated observations rather than corroborated investment signals.
Finally, AI capacity is increasingly a systems-engineering and organizational problem, not merely a chip problem. AMD faces first-launch latency risk at framework scale; conventional ROCm builds increase build time and artifact size, while tests show an approximately 74 millisecond SPIR-V JIT increment and approximately 407–414 millisecond fat-binary cold starts 68. Release-mode compile time rises roughly 340 milliseconds per added offload architecture 68. AMD’s architecture roadmap extends through at least 2028 37, and the company’s recovery was associated with the 2017 EPYC launch 41. Similar deployment constraints appear in automotive validation, Toyota’s possible two- to three-year delay from chip fabrication constraints, and long regulatory approval cycles 56,57. These examples reinforce that Alphabet’s advantage depends not only on model quality and hardware scale, but also on software integration, developer tooling, deployment reliability, and the ability to convert infrastructure into usable customer outcomes.
Implications for Alphabet
The principal conclusion is conditional. Under current conditions, HBM scarcity is both a competitive moat and a capital-allocation hazard. Alphabet’s integrated position should allow it to extract more value from each unit of scarce memory and compute than a less integrated competitor. The performance improvements reported for AlloyDB 39 and the broader move toward managed memory and efficient scheduling 38 support the view that optimization can raise effective capacity without proportionate hardware additions.
We must nevertheless distinguish temporary quasi-rents from durable market power. HBM is not proprietary, several companies can produce it, and future fab expansion or Chinese progress could pressure pricing 46. Memory markets outside HBM remain cyclical and commoditized 45. Historical hardware cycles show that supply gluts can destroy margins and prolong equity-market recoveries 48,51. Alphabet’s durable advantage is therefore more likely to arise from scale, procurement, utilization, and software differentiation than from ownership of scarce components alone.
The immediate strategic implication is to secure long-dated capacity while developing architectures that reduce dependence on the most expensive memory tiers. HBM should remain critical for training and high-performance inference, but HBF, DRAM, compressed databases, retrieval systems, selective checkpointing, and persistent memory-management layers may improve the economics of less latency-sensitive workloads 39,66. The central uncertainty is the relative speed of two opposing forces: efficiency gains may reduce hardware required per query, while model capability and expanding usage may increase total demand faster than efficiency improves. The resulting range of outcomes affects Google Cloud growth, capital expenditure, depreciation expense, and operating margins.
Indicators for investors
Investors should monitor four practical indicators:
- Commitment discipline: the duration and pricing of Alphabet’s advance memory and accelerator commitments.
- Utilization: the rate at which Google converts hardware into billable Google Cloud usage.
- Software efficiency: evidence that model and systems improvements are reducing inference cost per query.
- Supply normalization: signs that HBM capacity expansion is beginning to moderate pricing and supplier margins.
The evidence supports a constructive long-term view of Alphabet’s infrastructure position, but it does not support treating current AI demand or memory scarcity as a permanent, linear growth driver. The more defensible conclusion is that Alphabet is well placed to navigate a period of constrained supply, provided it continues to combine advance procurement with disciplined capacity planning and software-led utilization gains.
Key takeaways
- AI infrastructure—particularly HBM, DRAM, packaging, thermal management, and accelerator scheduling—remains the dominant constraint, with the most corroborated claims pointing toward supply shortages extending to approximately 2028 21,33,45,46.
- Alphabet’s strategic advantage is likely to arise from vertically integrated infrastructure and software efficiency, including database acceleration, memory management, custom accelerators, and cloud distribution 38,39.
- HBM scarcity supports near-term supplier and infrastructure economics, but standardization, new capacity, Chinese catch-up, and historical memory-cycle risk limit the case for permanent scarcity rents 46.
- The principal investment risk is overbuilding expensive infrastructure that becomes economically obsolete before its accounting life ends. Capacity commitments, utilization, depreciation, and inference-cost reductions therefore deserve closer attention than headline AI demand alone 19,23,26,59.