Analysis · deep cards

Nobody buys GPUs.

Every buyer purchases a unit of consumption matched to their team's level of abstraction. The AI infrastructure market is a ladder of those units, from megawatts to minutes of conversation. The companies below map the rungs, including rungs contested by opposed postures. Comparing their hourly H100 price is a category error; most of them don't primarily sell hours.

All public facts verified as of 4 July 2026 · by Rahul Patwardhan

AI minutes · voicewhole conversations, $ per minute all-in
GPU-secondsserverless time: scale-to-zero, no commitment
Workflow + hoursthe platform, plus hours on any fleet behind it
GPU-hoursreserved time on hardware the buyer operates
Clusters + sovereigntydedicated clusters, in-country by design
Megawattspower + whole AI factories
↑ most abstract: buyers never see hardwaremost physical: buyers sign for power ↓
The second axis

Every seller here is also a buyer

No spec sheet publishes what each of these companies buys. Runpod buys capacity from data-center partners and vetted hosts; it deployed three times its previous total fleet in Q1 2026 alone. DeepInfra buys data centers outright. AI Fabrik buys grid access before anything else, 150 MW under contract through GridCARE with connection in months rather than years, then fills modular edge sites with B300s. Lightning AI routes demand into seven or more third-party clouds while operating its own fleet. Firmus buys renewable power in hundred-megawatt blocks. Radiant goes one layer deeper. It buys permit-stabilized land banks, builds behind-the-meter generation, and racks NVIDIA systems on top, which makes it a power generator rather than a reseller. Vast.ai sits at the opposite limit: it owns no GPUs and buys no capacity at all. Over 1,400 independent providers list hardware they already paid for, and Vast now sells them the operating system: sourcing, financing against platform earnings, presold enterprise demand. Telnyx buys GPUs outright, racks them inside its own carrier PoPs, and buys no cloud capacity.

Knowing what a company buys tells you who its natural counterparty is, and deals happen there. The matrix below shows both sides.

The comparison

Both sides of the sheet

Company Unit sold Unit bought Silicon Scale signal Center of gravity Contract shape Natural buyer The bet
TelnyxAI minutes + tokensGPUs outright, carrier interconnectsOwned fleet (models undisclosed)4,000+ owned GPUs; ~$2.2M ever raised18 PoPs: US, EU, APACPer-minute ($0.08) · per-token, no minimumsVoice-AI & agent buildersThe carrier edge beats the cloud edge
FriendliAITokens, custom modelsGPU capacity across cloudsNVIDIA via partners (B300)6–7× rev growth; 25–30 enterprise clientsUS · Korea / APACPer-token · dedicated · BYO-clusterEnterprises with custom modelsInference is a software problem
DeepInfraTokensData centers, GPUsNVIDIA Blackwell, Rubin~5T tokens/week8 US data centersPer-token · DeepCluster 3–5yrApp builders / agentsAgentic inference eats compute
AI FabrikTokens, delivered at the edgeGrid access (150 MW via GridCARE), B300s, modular sitesNVIDIA B300, thousands deployingFirst of 5 sites live Jul 2026; incubated in Gruve; Mayfield-backedUS tier-1/tier-2 metro edgeUndisclosed · enterprise + model cosLatency-sensitive production inferenceTokens follow the CDN path
RunpodGPU-secondsPartner + host capacityNVIDIA + consumer~$240M ARR; 500K devs31 global regionsPer-second, no commitmentLong-tail devs & startupsLong tail > whales
Lightning AIWorkflow + multi-cloud hoursDemand routed to 7+ cloudsMulti: fleet H100/B200/GB30035K+ GPU fleet post-mergerUS fleet, global marketplaceSeats + on-demand + reservedEnterprise ML teamsWorkflow > any single fleet
TensorWaveAMD GPU-hoursAMD silicon, 2GW+ capacityAMD only, MI300X→MI355X$493M raised; 8,192 MI325X liveNorth AmericaReserved clustersMemory-bound inference platformsMemory/$ wins inference
Vast.aiGPU-hours, marketplaceNothing; supply is listed rather than boughtHost-owned, RTX 4090 → H20017K+ GPUs · 1,400+ providers (Feb 2026); 98.6% B200 utilization (Jul 2026); ~$4M ever raised500+ locations, the global long tailOn-demand · interruptible auction · reserved, per-second billingCost-sensitive researchers & batch jobsLiquidity beats ownership
GMI CloudClusters + sovereigntyPower, DCs, hardwareNVIDIA Blackwell → Rubin$12B Japan; ~7K GB300 TaiwanTaiwan–Japan–USReserved + on-demandAPAC inference & enterpriseSovereignty > spot market
CorvexClusters, confidentialNVIDIA allocation (NCP), DC capacity & powerNVIDIA H200/B200/GB200 → RubinNasdaq: MOVE (Mar 2026); 1st verified CC on HGX B200US Tier III+ DCs · hybrid + on-premMulti-year reserved · hybrid EKS nodesModel builders & regulated enterprisesSovereignty by cryptography rather than geography
FirmusMW / AI factoriesPower, land, GB300sNVIDIA GB300 → Rubin$10B debt + $1.35B equity; 1.6GW targetAustralia / APACMulti-year MW dealsHyperscaler / sovereignGreen tokens earn a premium
RadiantAI factories + GPU cloudPowered land, on-site generation, NVIDIA systemsNVIDIA Blackwell → Rubin (DSX design)$100B BAIIF pipeline; 5GW live / 45GW access claimed (Feb 2026)Global land bank · sovereign-firstLong-term contracts + Ori cloud on-demandSovereigns, telcos, select enterprisesCompute is a utility; capital cost wins
Private ledger : five fields tracked per provider, never published. They change weekly; we verify them directly with each provider.
Available capacity, next 90 daysGPU type · quantity · location · online date
Effective price floorreserved · per-GPU-hour, at volume
Idle-inventory posturecurrent utilization pressure
Preferred deal structureattach terms · co-sell terms
Capacity-partnerships ownerthe person who signs

The public rows earn the map. The private rows form the ledger, shared case-by-case with each provider's consent.

Ask about a specific provider →
One card per posture

The deep cards

Rung 1 · MegawattsFirmus

$10B Blackstone-led debt, 1.6GW Project Southgate, and a bet that green tokens earn a pricing premium.

Read the card →
Rung 1 · MW, the utility modelRadiant

Brookfield's compute vehicle: the Ori merger (Feb 2026), a $100B BAIIF pipeline, and claimed access to 5GW live power: the AI-factory-as-utility bet. Draft, pending verification.

Read the card →
Rung 2 · Clusters + sovereigntyGMI Cloud

~7,000 GB300s in Taiwan at near-full utilization and a $12B sovereign initiative in Japan.

Read the card →
Rung 2 · Clusters, confidentialCorvex

The security-first GPU cloud, now public (Nasdaq: MOVE); first verified confidential computing on HGX B200, model weights invisible even to the host. Draft, pending verification.

Read the card →
Rung 3 · GPU-hours, AMDTensorWave

The all-AMD cloud with Fireworks AI and Luma AI in production: the ROCm objection, answered with evidence.

Read the card →
Rung 3 · GPU-hours, marketplaceVast.ai

The pure marketplace: 17K+ GPUs from 1,400+ providers in 500+ locations, on ~$4M ever raised; 98.6% B200 fleet utilization, and now selling hosts the OS too: sourcing, financing, presold demand. Provider-reviewed.

Read the card →
Rung 4 · Workflow + hoursLightning AI

The PyTorch-native platform that just became supply: 35,000+ GPUs after the Voltage Park merger.

Read the card →
Rung 5 · GPU-secondsRunpod

~$240M ARR on ~$22M raised before its June round, and the biggest capacity buyer of the six.

Read the card →
Rung 6 · TokensDeepInfra

~5T tokens a week, ~30% from agents: the clearest public signal that agentic workloads are infrastructure-scale.

Read the card →
Rung 6 · Tokens, custom modelsFriendliAI

The continuous-batching inventor's asset-light answer to DeepInfra: sells tokens from your model, buys GPUs across clouds. Draft, pending verification.

Read the card →
Rung 6 · Tokens, delivered at the edgeAI Fabrik

The Gruve spinout building a CDN for tokens: modular sites near US metros, 150 MW under contract via GridCARE, first of five live July 2026. Draft, pending verification.

Read the card →
Rung 7 · AI minutesTelnyx

The carrier that became an inference provider: 4,000+ owned GPUs inside 18 telephony PoPs, selling conversation minutes at $0.08 all-in, on ~$2.2M ever raised. Draft, pending verification.

Read the card →
Questions this page answers

Frequently asked

What is the compute ladder?

A classification of AI infrastructure providers by the unit of consumption they sell, not the hardware they run: megawatts → AI factories → GPU clusters → GPU-hours → GPU-seconds → tokens → AI minutes. Buyers should shop the rung that matches their team's abstraction level.

What's the difference between a GPU cloud and an inference provider?

A GPU cloud sells time on hardware the buyer must operate; an inference provider sells the output, tokens through an API, and hides the hardware. Companies now span rungs: DeepInfra sells tokens and dedicated clusters; Lightning sells a platform, its own fleet, and third-party capacity.

Which of these companies also buy capacity?

Runpod (from DC partners and vetted hosts), DeepInfra (data centers outright), and Lightning AI (demand routed to 7+ clouds). The "both-sided" companies are the most active counterparties in the market; they transact in two directions.

Why these companies, and why is the top of the ladder crowded?

We add companies rung by rung so the ladder spans every unit, and contested rungs show every posture: around tokens, DeepInfra owns the metal, FriendliAI buys capacity, AI Fabrik sells proximity from edge sites near US metros, and Telnyx owns the metal and the network, selling minutes above it; around clusters, GMI Cloud sells sovereignty by geography while Corvex sells it by cryptography; on GPU-hours, TensorWave owns an all-AMD fleet while Vast.ai owns no GPUs at all, a 1,400-provider auction; on megawatts, Firmus sells a renewable premium while Radiant sells Brookfield-scale capital and powered land. The full market lives on the ecosystem map; cards ship as we verify them.

Methodology

How these cards are built

Every public figure carries a date and traces to a primary source: company newsrooms, funding press releases, filings, and named-customer announcements. We re-verify cards on update and stamp them. The five private-ledger fields come directly from providers in conversation and stay unpublished; we share them case-by-case with the provider's consent. Spot an error or a stale number? Tell me; corrections ship within a week.