Tensor Cores take operands from registers and shared memory, compute D = A×B + C inside a dot-product engine, and write back; field note 01 argues they are register-fed matrix engines rather than systolic arrays, tuned for low latency, dynamic shapes, small GEMMs and attention kernels. Each generation drifts toward the dataflow camp while keeping the unified execution model: bigger shared SRAM, TMA async copy engines, warp specialization, compiler-driven tiling. Around 80% of the accelerator market runs here, so most providers on the map default to this column.
Field note 01 — are Tensor Cores really systolic arrays? ↓Every chip is a position on data movement
Open a pre-tensor-core GPU and count what the transistors do. Each multiply-accumulate sits behind a register file and an operand-selection network; multiply that across thousands of SIMD lanes, three operands per instruction, plus bypass paths and forwarding logic, and routing hardware outnumbers math hardware. A systolic array deletes the routing question. The compiler fixes the path in advance, operands flow from processing element to processing element, and the die area freed from routing turns into compute. The TOPS/W advantage comes out of that trade, not out of cheaper multiplication.
A large array charges for this in latency (fill + compute + drain), a bill only big GEMMs amortize. And a transformer is more than GEMMs: softmax, normalization, masking and quant/dequant run between the matrix multiplies and want vector hardware. Both camps have noticed. Each new GPU generation adds dataflow machinery, each dataflow chip adds vector units, and the contest moves up into the compiler stack and the memory orchestration layer. Field notes 01–03 make the full argument.
Large systolic arrays surrounded by SRAM. The array fetches an operand once and reuses it as it passes through neighboring PEs, so register-file bandwidth and wire energy collapse.
Vector processors with deep scheduling, caches and communication fabrics. Any operand to any ALU at any time; whatever shape the workload takes, the hardware absorbs it.
Eight architectures, one spectrum
Each card places a family on the spectrum, states what it optimizes for, and lists who on the map runs it. Deployment chips link to Compute Ladder deep cards where they exist.
CDNA Matrix Cores sit in the same register-fed camp as NVIDIA. The wedge is memory per dollar: more HBM per GPU fits larger models on one node and cuts nodes per deployment, so memory-bound inference became the beachhead. TensorWave answers the ROCm objection with production customers, Fireworks AI among them, instead of benchmark decks.
The defining chip of the move-data-least pole. Operands propagate neighbor-to-neighbor through the MXU, partial sums flow through the PE network, and reuse collapses register-file bandwidth and wire energy. The design buys throughput and FLOPs/W on large, well-shaped workloads and leans on the XLA compiler to make everything else fit the pipe. Google sells it only through Google Cloud, at frontier-lab scale.
Amazon built its own answer to the TPU thesis and sells it inside AWS as a price-performance lever, never as merchant silicon. The NeuronCore already looks like the converged design: systolic tensor engines for the GEMMs, vector and scalar engines beside them for the rest of the transformer. Anchor customers fund the flywheel; Anthropic's Project Rainier cluster is the visible one.
Intel's thesis is networking: put standard Ethernet on the die and scale clusters on open fabrics instead of proprietary interconnects. Gaudi 3 competes on cost per token for mainstream inference and fine-tuning. Buyers' main hesitation is roadmap continuity while Intel folds its accelerator line into Falcon Shores. Deployment list draft, pending verification.
No HBM, no caches, no dynamic scheduling. The compiler places every operand movement before the program runs, which fixes latency at compile time and produces the fastest per-user token rates on the market. Serving a model means sharding its weights across many chips of SRAM, so the economics close only at high utilization on a known model set. Groq sells no chips into clouds; the cloud is the product.
Cerebras pushes move-data-least to its physical limit: keep the whole array on one wafer, and weights and activations never cross a package boundary. The company sells systems and, more and more, tokens; Cerebras Inference competes with GPU providers on tokens per second per user.
The hedge between the poles. The RDU reconfigures its dataflow spatially per model graph, which keeps dataflow efficiency while tolerating more shapes than a fixed array. Three memory tiers let one system serve many models and long contexts. Like Groq and Cerebras, SambaNova takes the silicon to market as an inference cloud.
All roads route through HBM
Register-fed architectures live and die by HBM bandwidth; the SRAM-heavy designs above exist to avoid buying it. Three companies make the part everyone else is waiting on.
Who runs what
The Compute Ladder sorts providers by the unit they sell. This matrix sorts them by the silicon underneath, the decision that sets their memory economics, their software surface, and their exposure to one vendor's roadmap. Rows marked draft await provider verification.
| Provider | Silicon fleet — today → next | Silicon posture | The silicon bet |
|---|---|---|---|
| Hyperscalers multi-silicon by design — NVIDIA plus in-house accelerators | |||
| AWS | NVIDIA H100/H200/Blackwell + Trainium2 → Trainium3 + Inferentia | Dual-track — merchant + in-house systolic | Own the price-performance floor with captive silicon; keep NVIDIA for demand it can't move |
| Microsoft Azure | NVIDIA H100/H200/GB200 + AMD MI300X + Maia in-house Maia scale — draft | Triple-source | Never negotiate with one vendor again |
| Google Cloud | TPU v6e → v8 + NVIDIA A3/A4 | Dual-track — in-house systolic first | The only hyperscaler whose in-house chip trains frontier models |
| Oracle Cloud | NVIDIA bare-metal superclusters + AMD MI300X | Dual-source, NVIDIA-heavy | RDMA-native bare metal beats virtualized fleets |
| GPU clouds silicon posture is the strategy: single-vendor depth to owning nothing | |||
| CoreWeave | NVIDIA H100/H200 → GB200/GB300 | NVIDIA-only, first-to-new-silicon | Be NVIDIA's fastest deployment arm and the allocation follows |
| Lambda | NVIDIA H100/H200 → Blackwell | NVIDIA-only | 1-Click Clusters — make NVIDIA silicon self-serve |
| Nebius | NVIDIA Hopper → Blackwell | NVIDIA-only, full-stack | EU full-stack AI cloud on the default silicon |
| TensorWave | AMD only — MI300X → MI355X, 8,192 MI325X live | Single-vendor challenger | Memory/$ wins inference; be the AMD cloud everyone benchmarks on |
| Vast.ai | Host-owned — RTX 4090 → H200, 17K+ GPUs listed | Owns nothing — marketplace | Liquidity beats ownership; the long tail prices the floor |
| Runpod | NVIDIA + consumer via partners and vetted hosts | Buys capacity, silicon-agnostic-ish | The unit is GPU-seconds; the silicon is whatever serves them cheapest |
| Lightning AI | Multi — H100/B200/GB300, 35K+ fleet + 7+ clouds routed | Multi-fleet router | Workflow > any single fleet — silicon is behind the abstraction |
| GMI Cloud | NVIDIA Blackwell → Rubin, ~7K GB300 Taiwan | NVIDIA-only, sovereign | Sovereignty by geography, on reference silicon |
| Corvex | NVIDIA H200/B200/GB200 → Rubin | NVIDIA-only, confidential-compute-first | Sovereignty by cryptography — weights invisible even to the host |
| Firmus | NVIDIA GB300 → Rubin | NVIDIA-only, renewable-powered | Green tokens earn a premium on the same silicon |
| Radiant | NVIDIA Blackwell → Rubin (DSX reference design) | NVIDIA-only, utility model | Silicon is a pass-through; capital cost of power wins |
| Inference providers the buyer never sees the hardware; the hardware still sets the token price | |||
| DeepInfra | NVIDIA Blackwell → Rubin, owned in 8 US DCs | NVIDIA-only, owns the metal | Own the depreciation curve, win the token price war |
| FriendliAI | NVIDIA via partners (B300) | Buys capacity, asset-light | Inference is a software problem — silicon is rented |
| Together AI | NVIDIA — owned GPU cloud + inference draft | NVIDIA-only, vertically integrated | Research-grade kernels on default silicon |
| Fireworks AI | NVIDIA + AMD MI300X (production on TensorWave) | Multi-silicon — rare among inference providers | Route each model to the silicon where its economics close |
| Baseten | NVIDIA across multiple clouds draft | Buys capacity, multi-cloud | Inference infra as product; silicon abstracted |
| Telnyx | Owned fleet — models undisclosed, 4,000+ GPUs in 18 PoPs | Owns the metal + the network | The carrier edge beats the cloud edge, whatever the SKU |
| Chip-owned clouds the silicon is the company, sold as tokens | |||
| Groq | LPU — SRAM-only, deterministic | Own silicon, own cloud | Latency you can compute, not measure |
| Cerebras | WSE — wafer-scale | Own silicon — systems + tokens | The cluster is one die; the interconnect problem dissolves |
| SambaNova | RDU — reconfigurable dataflow | Own silicon — systems + tokens | Dataflow efficiency without a fixed pipeline |
Rows trace to the map, the Compute Ladder deep cards, and public disclosures. Silicon fleets change quarterly; corrections ship within a week.
Correct a row →The architecture behind the choices
The sections above state deployment facts. The three notes below argue architecture. A contributor wrote them, and each carries an attribution and date instead of the map's verification stamp.
Are NVIDIA Tensor Cores really systolic arrays?
I increasingly think the answer is: not in the classical TPU sense. A more accurate mental model is that Tensor Cores are register-fed matrix engines composed of many dot-product units. The distinction may sound semantic, but it profoundly impacts energy efficiency, latency, compiler design, and workload suitability.
A classical systolic array is defined by its dataflow:
- Neighbor-to-neighbor communication
- Operands physically propagating through the array
- Partial sums flowing through the PE network
- Massive operand reuse
This dramatically reduces register-file bandwidth and wire energy, because an operand can be fetched once and reused many times as it moves through neighboring PEs. Tensor Cores appear fundamentally different. From the programmer's perspective:
D = A × B + CInputs come from registers and shared memory, computation happens inside the matrix engine, and results are written back to registers. There is no exposed programming model where activations and weights march across a large PE mesh.
So why didn't NVIDIA simply build a giant systolic array inside a GPU? Because GPUs optimize for low latency, dynamic tensor shapes, small GEMMs, attention kernels, and diverse workloads — and large systolic arrays pay a hidden tax:
Latency ≈ fill + compute + drainFor very large GEMMs, the fill/drain overhead is amortized. But for many GPU workloads — small matrices, inference, dynamic shapes — the array may spend a significant fraction of its time filling and draining instead of computing.
Open question: has the term "systolic array" become too loosely applied in our industry?
Why systolic arrays win: the real bottleneck was never compute
For a long time, matrix multiplication looked like a compute problem. It really wasn't. The limiting factor was always data movement, not arithmetic.
In pre-tensor-core GPUs, every MAC effectively sits behind a register file and a large operand-selection network. Now scale that across thousands of SIMD lanes, with three operands per instruction, multiported register files, bypass paths, and forwarding logic — the picture changes quickly. The ALU is still doing a simple multiply–accumulate, but most of the chip is now dedicated to getting the right operands to the right place at the right time.
At scale, the problem stops being "how do we compute faster" and becomes "how do we route data without collapsing under wiring and control complexity." Most of the silicon is not doing math. It is deciding where data should come from and how it should reach the ALU without conflicts.
Systolic arrays change this assumption entirely. Instead of pulling operands from a global structure every cycle, they remove the global structure from the critical path. Data is not selected and routed on demand — it is pushed through a fixed spatial pipeline. Each processing element does one thing: multiply, accumulate, and forward data locally. There is no per-cycle global register file access, no wide multiplexing tree, no complex routing decision for every operand.
Once this happens, the structure of the chip changes fundamentally. The hardware that previously existed to manage data movement largely disappears: wire networks shrink, control logic simplifies, and the area that was spent on register files and operand routing becomes available for actual compute units.
This is where the efficiency comes from. Not because multiplication itself becomes cheaper, but because the system stops spending most of its energy and area on moving data around.
That shift is what turns matrix multiplication from a routing problem back into a compute problem — and that is why systolic arrays deliver higher TOPS/W.
GPUs and TPUs are converging on the same middle ground
The AI hardware world is slowly converging toward a fascinating middle ground between GPUs and TPUs. TPUs were designed around a simple idea — move data as little as possible: large systolic arrays surrounded by SRAM achieve incredible FLOPs/W by maximizing data reuse and minimizing expensive memory movement. GPUs evolved differently — flexible vector processors with sophisticated scheduling, caches, and communication fabrics.
But modern AI workloads exposed an important reality: data movement energy dominates compute energy. That's why every new GPU generation is becoming more "dataflow-oriented":
- Tensor Cores — small matrix engines inside the SM
- Shared SRAM close to compute
- TMA / async copy engines
- Warp specialization
- Compiler-driven tiling and scheduling
Essentially, GPUs are adopting TPU ideas without sacrificing programmability. At the same time, TPUs face their own challenge: modern models are not just GEMMs. Transformers require softmax, normalization, masking, routing, quant/dequant, activations, sparse ops — creating communication overhead between vector units and matrix engines. GPUs still have an advantage here, because vector ALUs and tensor cores are tightly integrated within a unified execution model.
The result? The industry is converging toward hybrid architectures:
- Local SRAM
- Matrix engines
- Vector units
- Explicit async data movement
- Compiler-orchestrated scheduling
Increasingly, the real differentiator is no longer just hardware. It is the compiler stack and the memory orchestration layer. The future belongs to architectures that can minimize data movement without losing flexibility.
Frequently asked
What chips do AI clouds actually run?
NVIDIA, overwhelmingly: Hopper fleets moving to Blackwell, Rubin next. The exceptions that matter: TensorWave's all-AMD fleet (Fireworks AI in production), AMD instances at Azure and Oracle, TPUs at Google Cloud, Trainium at AWS, and Groq, Cerebras and SambaNova serving tokens on their own silicon.
Are Tensor Cores systolic arrays?
Not in the classical sense. They are register-fed matrix engines built from dot-product units, with no exposed PE mesh for operands to march across. The distinction shows up in energy, latency and compiler design; field note 01 makes the full argument.
Why does the silicon choice matter if I'm buying tokens?
Because the silicon sets the provider's memory economics and therefore your token price floor and latency profile. SRAM-only architectures (Groq) buy latency with capacity constraints; high-HBM chips (MI300X-class) serve large models on fewer nodes; NVIDIA buys flexibility and ecosystem at the market-clearing price.
What are field notes, and how are they different from the rest of the page?
The spectrum, family cards and matrix carry deployment facts, sourced and dated like the rest of the map. Field notes carry contributed architecture analysis. Each one names its author and date, sits on dark ground, and states a view the map does not independently verify.
How this page is built
This page keeps two kinds of content separate. Deployment facts (who runs what silicon) trace to the map, the Compute Ladder deep cards, and public disclosures: product pages, filings, named-deployment announcements. A row carries the draft flag until the provider confirms it. Field notes carry contributed analysis under the author's name and date. Spot a stale fleet or a wrong SKU? Tell me; corrections ship within a week.