Analysis · the chips layer, cut open

Compute was never the bottleneck.

Matrix multiplication looked like a compute problem. Data movement was the real limit. Each accelerator below answers one question: how far should an operand travel per FLOP? The answers form a spectrum. Every provider on the map has picked a point on it, and the chip designers are converging on the middle.

Deployment facts drawn from the map & public disclosures as of 16 July 2026 · field notes contributed, see attribution · by Rahul Patwardhan

Move the data as little as possible Make the data available as quickly as possible
Deterministic dataflow Operands march through fixed spatial pipelines: systolic arrays and their descendants. The design buys throughput and TOPS/W and pays in fill/drain latency and shape flexibility.
The converging middle Local SRAM + matrix engines + vector units + explicit async data movement + compiler-orchestrated scheduling. Each camp's newest designs land here.
← GPUs drift this way every generation · dataflow chips add vector units →
Register-fed matrix engines D = A×B+C from registers and shared memory; no exposed PE mesh. Low latency, dynamic shapes, small GEMMs, attention kernels, everything else.
← throughput & TOPS/W · the workload must fit the pipelatency & flexibility · the pipe adapts to the workload →
The two philosophies

Every chip is a position on data movement

Open a pre-tensor-core GPU and count what the transistors do. Each multiply-accumulate sits behind a register file and an operand-selection network; multiply that across thousands of SIMD lanes, three operands per instruction, plus bypass paths and forwarding logic, and routing hardware outnumbers math hardware. A systolic array deletes the routing question. The compiler fixes the path in advance, operands flow from processing element to processing element, and the die area freed from routing turns into compute. The TOPS/W advantage comes out of that trade, not out of cheaper multiplication.

A large array charges for this in latency (fill + compute + drain), a bill only big GEMMs amortize. And a transformer is more than GEMMs: softmax, normalization, masking and quant/dequant run between the matrix multiplies and want vector hardware. Both camps have noticed. Each new GPU generation adds dataflow machinery, each dataflow chip adds vector units, and the contest moves up into the compiler stack and the memory orchestration layer. Field notes 01–03 make the full argument.

TPU philosophy "Move the data as little as possible."

Large systolic arrays surrounded by SRAM. The array fetches an operand once and reuses it as it passes through neighboring PEs, so register-file bandwidth and wire energy collapse.

GPU philosophy "Make the data available as quickly as possible."

Vector processors with deep scheduling, caches and communication fabrics. Any operand to any ALU at any time; whatever shape the workload takes, the hardware absorbs it.

The silicon families

Eight architectures, one spectrum

Each card places a family on the spectrum, states what it optimizes for, and lists who on the map runs it. Deployment chips link to Compute Ladder deep cards where they exist.

NVIDIARegister-fed matrix engines
Silicon Hopper H100/H200 → Blackwell B200/GB200/GB300 → Rubin  Memory HBM3E  Fabric NVLink + InfiniBand/Spectrum-X  Moat CUDA

Tensor Cores take operands from registers and shared memory, compute D = A×B + C inside a dot-product engine, and write back; field note 01 argues they are register-fed matrix engines rather than systolic arrays, tuned for low latency, dynamic shapes, small GEMMs and attention kernels. Each generation drifts toward the dataflow camp while keeping the unified execution model: bigger shared SRAM, TMA async copy engines, warp specialization, compiler-driven tiling. Around 80% of the accelerator market runs here, so most providers on the map default to this column.

Runs on the map
CoreWeaveLambdaNebiusCrusoe DeepInfraGMI CloudCorvexFirmusRadiant+ nearly everyone
Field note 01 — are Tensor Cores really systolic arrays? ↓
AMD InstinctRegister-fed · memory-first
Silicon MI300X/MI325X → MI355X  Memory HBM3E, class-leading capacity per GPU  Fabric Infinity Fabric + Ethernet  Stack ROCm

CDNA Matrix Cores sit in the same register-fed camp as NVIDIA. The wedge is memory per dollar: more HBM per GPU fits larger models on one node and cuts nodes per deployment, so memory-bound inference became the beachhead. TensorWave answers the ROCm objection with production customers, Fireworks AI among them, instead of benchmark decks.

Runs on the map
TensorWave — all-AMD fleetFireworks AI (via TensorWave)Microsoft AzureOracle Cloud
Google TPUClassical systolic array
Silicon TPU v6e → v8  Compute MXU systolic arrays + large local SRAM  Fabric ICI optical torus  Stack JAX / Pathways / XLA

The defining chip of the move-data-least pole. Operands propagate neighbor-to-neighbor through the MXU, partial sums flow through the PE network, and reuse collapses register-file bandwidth and wire energy. The design buys throughput and FLOPs/W on large, well-shaped workloads and leans on the XLA compiler to make everything else fit the pipe. Google sells it only through Google Cloud, at frontier-lab scale.

Runs on the map
Google Cloud (exclusive)Google DeepMindAnthropic — training + serving
Field note 02 — why systolic arrays win on TOPS/W ↓
AWS Trainium / InferentiaHyperscaler systolic
Silicon Trainium2 → Trainium3  Compute NeuronCore tensor engines (systolic) + vector/scalar engines  Fabric NeuronLink + EFA  Stack Neuron SDK

Amazon built its own answer to the TPU thesis and sells it inside AWS as a price-performance lever, never as merchant silicon. The NeuronCore already looks like the converged design: systolic tensor engines for the GEMMs, vector and scalar engines beside them for the rest of the transformer. Anchor customers fund the flywheel; Anthropic's Project Rainier cluster is the visible one.

Runs on the map
AWS (exclusive)Anthropic — Project Rainier
Intel GaudiLarge matrix engines · Ethernet-native
Silicon Gaudi 3  Compute large MMEs + programmable TPCs  Fabric RoCE Ethernet on-chip — no proprietary interconnect  Stack SynapseAI

Intel's thesis is networking: put standard Ethernet on the die and scale clusters on open fabrics instead of proprietary interconnects. Gaudi 3 competes on cost per token for mainstream inference and fine-tuning. Buyers' main hesitation is roadmap continuity while Intel folds its accelerator line into Falcon Shores. Deployment list draft, pending verification.

Runs on the map
IBM CloudIntel Tiber AI Cloud
Groq LPUDeterministic dataflow · compiler-scheduled
Silicon LPU  Memory on-chip SRAM only — no HBM  Scheduling statically compiled to the cycle, no dynamic arbitration  Sold as tokens via GroqCloud

No HBM, no caches, no dynamic scheduling. The compiler places every operand movement before the program runs, which fixes latency at compile time and produces the fastest per-user token rates on the market. Serving a model means sharding its weights across many chips of SRAM, so the economics close only at high utilization on a known model set. Groq sells no chips into clouds; the cloud is the product.

Runs on the map
GroqCloud (own silicon)Humain — Saudi sovereign build draft
Cerebras WSEWafer-scale dataflow
Silicon Wafer-Scale Engine  Memory tens of GB of on-wafer SRAM at PE-adjacent distance  Thesis the interconnect problem dissolves if the cluster is one die

Cerebras pushes move-data-least to its physical limit: keep the whole array on one wafer, and weights and activations never cross a package boundary. The company sells systems and, more and more, tokens; Cerebras Inference competes with GPU providers on tokens per second per user.

Runs on the map
Cerebras Cloud (own silicon)Core42 / G42 — Condor Galaxy
SambaNova RDUReconfigurable dataflow
Silicon RDU (Reconfigurable Dataflow Unit)  Compute spatially reconfigured per model graph  Memory large three-tier (SRAM/HBM/DDR)  Sold as tokens + dedicated systems

The hedge between the poles. The RDU reconfigures its dataflow spatially per model graph, which keeps dataflow efficiency while tolerating more shapes than a fixed array. Three memory tiers let one system serve many models and long contexts. Like Groq and Cerebras, SambaNova takes the silicon to market as an inference cloud.

Runs on the map
SambaNova Cloud (own silicon)
Field note 03 — the convergence ↓
The layer beneath the layer

All roads route through HBM

Register-fed architectures live and die by HBM bandwidth; the SRAM-heavy designs above exist to avoid buying it. Three companies make the part everyone else is waiting on.

HBM — who makes the memory every GPU is waiting on
SK HynixHBM3E leader with roughly 95% of GPU-attached memory.Leader · primary NVIDIA supplier
SamsungHBM3E plus CXL memory expansion for capacity beyond the package.Major · qualifying across vendors
MicronHBM3E and CXL pooling, the third source every buyer wants to exist.Major · US-based supply
The deployment matrix

Who runs what

The Compute Ladder sorts providers by the unit they sell. This matrix sorts them by the silicon underneath, the decision that sets their memory economics, their software surface, and their exposure to one vendor's roadmap. Rows marked draft await provider verification.

Provider Silicon fleet — today → next Silicon posture The silicon bet
Hyperscalers multi-silicon by design — NVIDIA plus in-house accelerators
AWSNVIDIA H100/H200/Blackwell + Trainium2 → Trainium3 + InferentiaDual-track — merchant + in-house systolicOwn the price-performance floor with captive silicon; keep NVIDIA for demand it can't move
Microsoft AzureNVIDIA H100/H200/GB200 + AMD MI300X + Maia in-house Maia scale — draftTriple-sourceNever negotiate with one vendor again
Google CloudTPU v6e → v8 + NVIDIA A3/A4Dual-track — in-house systolic firstThe only hyperscaler whose in-house chip trains frontier models
Oracle CloudNVIDIA bare-metal superclusters + AMD MI300XDual-source, NVIDIA-heavyRDMA-native bare metal beats virtualized fleets
GPU clouds silicon posture is the strategy: single-vendor depth to owning nothing
CoreWeaveNVIDIA H100/H200 → GB200/GB300NVIDIA-only, first-to-new-siliconBe NVIDIA's fastest deployment arm and the allocation follows
LambdaNVIDIA H100/H200 → BlackwellNVIDIA-only1-Click Clusters — make NVIDIA silicon self-serve
NebiusNVIDIA Hopper → BlackwellNVIDIA-only, full-stackEU full-stack AI cloud on the default silicon
TensorWaveAMD only — MI300X → MI355X, 8,192 MI325X liveSingle-vendor challengerMemory/$ wins inference; be the AMD cloud everyone benchmarks on
Vast.aiHost-owned — RTX 4090 → H200, 17K+ GPUs listedOwns nothing — marketplaceLiquidity beats ownership; the long tail prices the floor
RunpodNVIDIA + consumer via partners and vetted hostsBuys capacity, silicon-agnostic-ishThe unit is GPU-seconds; the silicon is whatever serves them cheapest
Lightning AIMulti — H100/B200/GB300, 35K+ fleet + 7+ clouds routedMulti-fleet routerWorkflow > any single fleet — silicon is behind the abstraction
GMI CloudNVIDIA Blackwell → Rubin, ~7K GB300 TaiwanNVIDIA-only, sovereignSovereignty by geography, on reference silicon
CorvexNVIDIA H200/B200/GB200 → RubinNVIDIA-only, confidential-compute-firstSovereignty by cryptography — weights invisible even to the host
FirmusNVIDIA GB300 → RubinNVIDIA-only, renewable-poweredGreen tokens earn a premium on the same silicon
RadiantNVIDIA Blackwell → Rubin (DSX reference design)NVIDIA-only, utility modelSilicon is a pass-through; capital cost of power wins
Inference providers the buyer never sees the hardware; the hardware still sets the token price
DeepInfraNVIDIA Blackwell → Rubin, owned in 8 US DCsNVIDIA-only, owns the metalOwn the depreciation curve, win the token price war
FriendliAINVIDIA via partners (B300)Buys capacity, asset-lightInference is a software problem — silicon is rented
Together AINVIDIA — owned GPU cloud + inference draftNVIDIA-only, vertically integratedResearch-grade kernels on default silicon
Fireworks AINVIDIA + AMD MI300X (production on TensorWave)Multi-silicon — rare among inference providersRoute each model to the silicon where its economics close
BasetenNVIDIA across multiple clouds draftBuys capacity, multi-cloudInference infra as product; silicon abstracted
TelnyxOwned fleet — models undisclosed, 4,000+ GPUs in 18 PoPsOwns the metal + the networkThe carrier edge beats the cloud edge, whatever the SKU
Chip-owned clouds the silicon is the company, sold as tokens
GroqLPU — SRAM-only, deterministicOwn silicon, own cloudLatency you can compute, not measure
CerebrasWSE — wafer-scaleOwn silicon — systems + tokensThe cluster is one die; the interconnect problem dissolves
SambaNovaRDU — reconfigurable dataflowOwn silicon — systems + tokensDataflow efficiency without a fixed pipeline

Rows trace to the map, the Compute Ladder deep cards, and public disclosures. Silicon fleets change quarterly; corrections ship within a week.

Correct a row →
Field notes · contributed analysis

The architecture behind the choices

The sections above state deployment facts. The three notes below argue architecture. A contributor wrote them, and each carries an attribution and date instead of the map's verification stamp.

Field note 01 · NVIDIA Tensor Cores attaches to: NVIDIA · Google TPU

Are NVIDIA Tensor Cores really systolic arrays?

I increasingly think the answer is: not in the classical TPU sense. A more accurate mental model is that Tensor Cores are register-fed matrix engines composed of many dot-product units. The distinction may sound semantic, but it profoundly impacts energy efficiency, latency, compiler design, and workload suitability.

A classical systolic array is defined by its dataflow:

  • Neighbor-to-neighbor communication
  • Operands physically propagating through the array
  • Partial sums flowing through the PE network
  • Massive operand reuse

This dramatically reduces register-file bandwidth and wire energy, because an operand can be fetched once and reused many times as it moves through neighboring PEs. Tensor Cores appear fundamentally different. From the programmer's perspective:

D = A × B + C

Inputs come from registers and shared memory, computation happens inside the matrix engine, and results are written back to registers. There is no exposed programming model where activations and weights march across a large PE mesh.

So why didn't NVIDIA simply build a giant systolic array inside a GPU? Because GPUs optimize for low latency, dynamic tensor shapes, small GEMMs, attention kernels, and diverse workloads — and large systolic arrays pay a hidden tax:

Latency ≈ fill + compute + drain

For very large GEMMs, the fill/drain overhead is amortized. But for many GPU workloads — small matrices, inference, dynamic shapes — the array may spend a significant fraction of its time filling and draining instead of computing.

TPU MXUClassical systolic array — optimized for throughput and efficiency.
Tensor CoresRegister-fed matrix engines — optimized for flexibility and low latency.

Open question: has the term "systolic array" become too loosely applied in our industry?

Contributed by [Contributor name — role]Compiled 16 July 2026
Field note 02 · Systolic arrays attaches to: TPU · Trainium · Groq

Why systolic arrays win: the real bottleneck was never compute

For a long time, matrix multiplication looked like a compute problem. It really wasn't. The limiting factor was always data movement, not arithmetic.

In pre-tensor-core GPUs, every MAC effectively sits behind a register file and a large operand-selection network. Now scale that across thousands of SIMD lanes, with three operands per instruction, multiported register files, bypass paths, and forwarding logic — the picture changes quickly. The ALU is still doing a simple multiply–accumulate, but most of the chip is now dedicated to getting the right operands to the right place at the right time.

At scale, the problem stops being "how do we compute faster" and becomes "how do we route data without collapsing under wiring and control complexity." Most of the silicon is not doing math. It is deciding where data should come from and how it should reach the ALU without conflicts.

Systolic arrays change this assumption entirely. Instead of pulling operands from a global structure every cycle, they remove the global structure from the critical path. Data is not selected and routed on demand — it is pushed through a fixed spatial pipeline. Each processing element does one thing: multiply, accumulate, and forward data locally. There is no per-cycle global register file access, no wide multiplexing tree, no complex routing decision for every operand.

Once this happens, the structure of the chip changes fundamentally. The hardware that previously existed to manage data movement largely disappears: wire networks shrink, control logic simplifies, and the area that was spent on register files and operand routing becomes available for actual compute units.

This is where the efficiency comes from. Not because multiplication itself becomes cheaper, but because the system stops spending most of its energy and area on moving data around.

Pre-tensor-core designsOptimized for flexibility: any operand to any ALU at any time.
Systolic arraysOptimized for locality: operands move in predictable patterns and compute happens where they land.

That shift is what turns matrix multiplication from a routing problem back into a compute problem — and that is why systolic arrays deliver higher TOPS/W.

Contributed by [Contributor name — role]Compiled 16 July 2026
Field note 03 · The convergence attaches to: the whole spectrum

GPUs and TPUs are converging on the same middle ground

The AI hardware world is slowly converging toward a fascinating middle ground between GPUs and TPUs. TPUs were designed around a simple idea — move data as little as possible: large systolic arrays surrounded by SRAM achieve incredible FLOPs/W by maximizing data reuse and minimizing expensive memory movement. GPUs evolved differently — flexible vector processors with sophisticated scheduling, caches, and communication fabrics.

But modern AI workloads exposed an important reality: data movement energy dominates compute energy. That's why every new GPU generation is becoming more "dataflow-oriented":

  • Tensor Cores — small matrix engines inside the SM
  • Shared SRAM close to compute
  • TMA / async copy engines
  • Warp specialization
  • Compiler-driven tiling and scheduling

Essentially, GPUs are adopting TPU ideas without sacrificing programmability. At the same time, TPUs face their own challenge: modern models are not just GEMMs. Transformers require softmax, normalization, masking, routing, quant/dequant, activations, sparse ops — creating communication overhead between vector units and matrix engines. GPUs still have an advantage here, because vector ALUs and tensor cores are tightly integrated within a unified execution model.

The result? The industry is converging toward hybrid architectures:

  • Local SRAM
  • Matrix engines
  • Vector units
  • Explicit async data movement
  • Compiler-orchestrated scheduling

Increasingly, the real differentiator is no longer just hardware. It is the compiler stack and the memory orchestration layer. The future belongs to architectures that can minimize data movement without losing flexibility.

Contributed by [Contributor name — role]Compiled 16 July 2026
Questions this page answers

Frequently asked

What chips do AI clouds actually run?

NVIDIA, overwhelmingly: Hopper fleets moving to Blackwell, Rubin next. The exceptions that matter: TensorWave's all-AMD fleet (Fireworks AI in production), AMD instances at Azure and Oracle, TPUs at Google Cloud, Trainium at AWS, and Groq, Cerebras and SambaNova serving tokens on their own silicon.

Are Tensor Cores systolic arrays?

Not in the classical sense. They are register-fed matrix engines built from dot-product units, with no exposed PE mesh for operands to march across. The distinction shows up in energy, latency and compiler design; field note 01 makes the full argument.

Why does the silicon choice matter if I'm buying tokens?

Because the silicon sets the provider's memory economics and therefore your token price floor and latency profile. SRAM-only architectures (Groq) buy latency with capacity constraints; high-HBM chips (MI300X-class) serve large models on fewer nodes; NVIDIA buys flexibility and ecosystem at the market-clearing price.

What are field notes, and how are they different from the rest of the page?

The spectrum, family cards and matrix carry deployment facts, sourced and dated like the rest of the map. Field notes carry contributed architecture analysis. Each one names its author and date, sits on dark ground, and states a view the map does not independently verify.

Methodology

How this page is built

This page keeps two kinds of content separate. Deployment facts (who runs what silicon) trace to the map, the Compute Ladder deep cards, and public disclosures: product pages, filings, named-deployment announcements. A row carries the draft flag until the provider confirms it. Field notes carry contributed analysis under the author's name and date. Spot a stale fleet or a wrong SKU? Tell me; corrections ship within a week.