Skip to content
Some content is members-only. Sign in to access.

TPU Bull Case: $51B by 2027 vs. Bear Case: Supplier Chokepoints and CUDA Lock-in

Weighing Piper Sandler's forecast against Broadcom leverage, Nvidia efficiency gaps, and unproven external demand

By KAPUALabs

Alphabet’s custom Tensor Processing Units now carry more strategic weight than any other single asset in its AI infrastructure stack. What began as an internal substitute for Nvidia hardware has become a full-stack platform—chip, model, deployment, and cloud—that Alphabet is now beginning to sell to outside customers 26,32,53. The test of the next two years is whether that integration proves more than an operational convenience: whether it becomes a durable revenue engine.

A Decade of In-House Capacity

TPUs are Google-designed application-specific integrated circuits specialized for AI workloads 17,26. Alphabet has been building them for about a decade 53, starting with internal search, advertising, and translation workloads around 2015 7,37. The original logic was defensive: Google used them in-house to avoid buying Nvidia chips and to reduce costs 6,55. Today the line has reached its seventh generation, Ironwood 6,37,55, and is embedded across Compute Engine machine families such as TPU7x, TPU v6e, and TPU v5p 3,17.

The product line is now segmented by workload. Trillium, the v6e, is recommended for training and fine-tuning Transformers and for large-scale inference with Gemma 2, Llama, and diffusion models 17; v5p is aimed at massive-scale multimodal training and large recommendation systems 17; Ironwood is the highest-performance option for dense models, mixture-of-experts models, intensive pre-training, and decode-heavy inference 17,18.

The scale and interconnect architecture are equally significant. A v5p Pod comprises 8,960 TPU chips interconnected by reconfigurable high-speed links 17; Ironwood scales to 9,216 chips per pod 17, while Trillium pods hold 256 chips 17. Inside a Pod or slice, chips communicate over dedicated optical links in 2D or 3D torus meshes that bypass traditional network stacks 13. Earlier generations used inter-chip interconnect within a slice and a single NIC for host traffic; Trillium and Ironwood introduce a native multi-NIC architecture 13, and Cloud TPU Multislice links independent meshes over Google’s Jupiter Data Center Network for models beyond one slice 13. This is not merely a chip-improvement story; it is a networking and pod-scaling story.

The Efficiency Ledger Has Both Entries

Google’s argument for TPUs is efficiency and total cost. The post claims a distinct inference-efficiency advantage for TPU V7 and a structural advantage in aggregate inference throughput 35, along with structural advantages in power profile and long-run total cost of ownership 35. It reports V7 at 5.42 TFLOPS/W for both FP32 training and FP8 inference, corresponding to 4.61 PFLOPS at 0.85 kW 35, and says V7 delivers more than twice the inference throughput per watt under lower power envelopes 35. Alphabet’s board-level cost case is that TPU v5p and Trillium reduce computing costs compared with relying only on third-party GPUs 43, that TPU pricing is roughly 40% below Nvidia equivalents 55, and that the custom TPU payback period is roughly half the average server payback period 50.

Those numbers are not uncontested. A separate aggregate comparison gives Google TPU an efficiency of 4.25 TFLOPS/W 35, and the post characterizes Nvidia’s GB200 as about twice as energy-efficient as TPU V7 for raw FP32 training 35. The honest reading is that Google’s advantage is strongest in inference and matrix-heavy workloads, not across every benchmark. The source’s own qualifications align: TPU economics are favorable when a workload benefits from Google’s pod scaling architecture 18, and capacity price is not the same as workload economics 18.

Software remains the hinge. Google is expanding native PyTorch and JAX support to reduce porting friction tied to developers’ dependence on Nvidia CUDA 35, and the integrated system is positioned as using Google’s software stack rather than Nvidia CUDA 37 with liquid-cooled pods 37. But the material also states plainly that CUDA-based dependence means many Google Cloud customers still want Nvidia GPUs 37, and Google continues to support Nvidia GPUs in its Cloud 5. The moat is being bridged, not yet closed.

Looking ahead, TPU 8t is designed for training and reportedly provides up to three times the processing power of Ironwood 5,21, while TPU 8i is designed for inference and reinforcement learning with a claimed 80% better performance per dollar 5,21. Both are described as coming soon for customers 5. These are promises, not yet delivering revenue.

From Rental to Revenue: The 2027 Test

Alphabet has long rented TPUs through Google Cloud but did not sell them 5,41. That changed when it began selling TPU systems to outside customers 6,26, recognizing external hardware sales for the first time in the second quarter of 2026 26,48. The systems include hardware, software, installation, and support 41, and Alphabet earns directly from them while retaining buyers as Google Cloud customers 41.

But scale remains prospective. Revenue from directly delivered TPU systems was described as small so far 41, and Piper Sandler’s projections—approximately $8 billion in 2026 and $51 billion in 2027 36—are explicitly forecasts, not reported results 36. Most TPU sales are expected to occur in 2027 2,47, and some analysts see substantial external revenue only by 2028 11. Cloud growth has already accelerated without the first system sales 56, which makes TPU’s contribution difficult to isolate 5.

Demand signals are present but soft. Citi attributes Google Cloud’s projected acceleration to demand for Google’s TPUs 42, and TPU-related cloud revenue is expected to accelerate as multi-billion-dollar customer contracts convert 4; reports also describe demand for Alphabet’s TPUs exceeding available manufacturing capacity 25. One concrete external commitment is an expanded collaboration involving Alphabet and Broadcom that provides Anthropic with access to roughly 3.5 gigawatts of TPU capacity beginning in 2027 5. Yet capacity access agreements may not translate cleanly into recognized revenue depending on ramp and whether TPUs meet expectations 5.

Pricing, for its part, is segmented rather than singular: stated per chip-hour and varying by region, purchasing model, number of chips, and required accelerator quality and service level 18.

A Supplier Base in Motion

Where Alphabet does not yet control its own destiny is below the chip architecture: Broadcom and TSMC. Broadcom has co-developed TPUs with Google for about ten years 7, and the two have agreements for development and supply of next-generation AI racks and TPUs through 2031 5,41. Google is described as Broadcom’s anchor XPU customer and as dependent on Broadcom 7,37, and it also depends on TSMC 37. That is a concentrated chokepoint for a company trying to industrialize a proprietary accelerator. Broadcom’s own AI semiconductor engine is vast—fiscal Q3 2026 AI semiconductor revenue reached $16.7 billion 53—a scale that secures supply but also gives the partner pricing and allocation leverage.

Reported plans to diversify are the clearest signal of pressure in that relationship. A post on X describes MediaTek as reportedly winning both design contracts for Google’s 2028 TPU generation, code-named ‘Humufish’ and ‘Triggerfish’ 28, and downstream material sketches a full-custom TPU V9 arrangement in which Google directly reserves HBM, advanced-packaging, and substrate capacity while sharing quotas with MediaTek 35. But these claims remain unverified supply-chain assertion without confirmation from Google 39; the V9 design has not taped out 35. Broadcom’s reported refusal of margin-dilutive semi-custom economics and its heavily committed pipeline are cited as reasons it was omitted from initial V9 planning 35, and Broadcom capacity is still described as reserved insurance, though limited 35. Tentative alternatives include a possible shift of the HBM4e inference SKU to Marvell because of customer hesitation about MediaTek silicon 35, but Broadcom and Marvell are not expected to operate under a full COT framework 35.

The binding constraint may not be design at all, but packaging and memory. Packaging capacity is identified as a supply-chain risk 46, and cumulative V8 shipments of 2.1 million units by year-end 2026 are conditional on alleviating HBM and CoWoS bottlenecks 35.

A Stress Test Above the Data Center Floor

Project Suncatcher is not a revenue program; it is a technical probe. Google is flying four Trillium TPUs into low Earth orbit aboard a SpaceX Falcon 9 8,12,31,32,40, with the mission designed to test how the chips withstand radiation, thermal extremes, and the physical stress of spaceflight while running a version of Google’s Gemma open-weight model 8,19. The payload is explicitly not a working orbital data center 29, and the idea of beginning orbital data centers by 2028 appears only as a thread-level aspiration 45. In other words, Google is hardening its hardware and gathering data, not claiming a second cloud region in orbit.

On the ground, Google’s Trillium TPUs were exposed to a proton beam while running AI workloads at UC Davis’s Crocker Nuclear Laboratory, and they ‘hold up remarkably well’ according to Google 14,15,19,32. Four TPUs produce roughly 1 kW of heat 32, a modest but real thermal challenge for a solar-powered satellite 10.

The Cloud Context: Managing Scarcity, Then Buying Power

TPUs will also be judged against a GPU rental market that is rationed, contradictory, and likely to shift. Google retires old silicon aggressively: Nvidia P100 GPUs hit end of support on September 15, 2026 9,22, leaving L4 and RTX PRO 6000 as suggested replacements 22. New GPU capacity is controlled by global quotas, model-specific region quotas, and reservations that provide high assurance of availability 20.

Pricing signals are split. H100 rental prices are increasing 1,54, B200 rental price reportedly rose 33% in 2026 54, and a cited spot rate places a B200 at $5.78 per hour 44. But H100 and A100 rental rates were also described as flat to down 44,51, likely because advertised low rates rely on older or lower-memory tiers 49. That is not a single market price; it is a utilization and contract mix.

The more consequential risk for Google Cloud is a 2027 loosening of GPU supply, which could shift pricing power to buyers 34, and compute-rental margins depend directly on that balance 34. ASIC shipments may overtake Nvidia GPU shipments in 2027 for the first time 27,52, a shift that would reward a proprietary chip strategy if buyer attention moves from GPU count to tokens per megawatt 38.

That shift is happening under severe physical and component-cost pressure. An H100 can draw 700–1,200 W 23, liquid cooling is increasingly necessary for high-density GPU deployments 16, and Morgan Stanley estimates that over half of GPU servers sold from 2026 through 2028 may lack a power hookup 24,44,51. On the component side, UBS estimates HBM average pricing will rise about 79% in 2027 33, and memory/storage prices have spiked: average prices for 32GB DDR5 kits are up 363% since September 2025, and 2TB NVMe SSDs up 137% 24. That compresses the spread between rental revenue and infrastructure cost.

The Next Three Years Will Separate Integration from Revenue

Taken together, these facts point to a strategy of vertical leverage rather than outright substitution. Alphabet is not trying to displace Nvidia everywhere; it is trying to own the most efficient path for its own models, Gemini 30,37, and to offer that path as a differentiated Cloud product. The efficient-inference claims, the pod scaling, the multi-NIC architecture, and the orbital testing all serve the same objective: lower the lifetime cost of running AI workloads on Google infrastructure. That is the modern equivalent of building the mill alongside the railroad.

Three conditions will determine whether the strategy matures. First, 2027 must convert capacity and partner execution into recognized revenue; the most-cited figures are analyst forecasts, and the first external sales were small 36,41. Second, supply-chain diversification through MediaTek and Marvell must not trade a Broadcom dependency for a packaging or yield dependency; the roadmap as reported has not taped out, and the extra partners carry their own qualification risks 35,46. Third, buyer preference must shift from headline GPU availability to workload economics. If ASIC volumes rise and power remains the binding constraint, Alphabet’s TPU-to-Gemini stack is well positioned; if Nvidia’s GB200 continues to win raw training efficiency and CUDA lock-in persists, TPU growth will remain narrower than the forecasts suggest 35,37.

The industrial pattern is familiar: the producer that controls the cheapest integrated route to a scarce resource—here, compute per watt—does not need to win every benchmark. But until external sales scale and the 2027 supply picture resolves, TPUs are a proven internal advantage and an unproven revenue engine, not yet a modern trust in silicon.

More from KAPUALabs

See all
| Free

Alphabet Bull Case Meets YouTube Attribution Risk

By KAPUALabs
/
| Free

Alphabet Bull vs Bear: Cloud Margins, Capex, and Recurring EPS Outlook

By KAPUALabs
/
| Free

Alphabet: Durable Search Moat vs. Cloud Backlog Execution Risk

By KAPUALabs
/