Rung 6 · Tokens — delivered at the edge

AI Fabrik

Sells delivered inference: a CDN for tokens, built from modular token factories near US tier-1 and tier-2 cities, engineered around performance per watt and per dollar.

Draft, pending provider verification · compiled 18 July 2026  ·  official site: aifabrik.com  ·  incubated inside Gruve

Origin
Incubated inside Gruve, spun out 2026
Center of gravity
US tier-1 & tier-2 metro edge
Silicon
NVIDIA B300, thousands deploying
Backers
Mayfield · Xora · Acclimate · Cisco Investments
The two-sided view

What AI Fabrik sells, and what it buys

Atomic unit sold

Delivered inference. The company runs an inference delivery network, which works the way a CDN works for content: modular token factories placed near US tier-1 and tier-2 cities serve model responses close to users, with latency, cost and compliance handled at the site level. Each site is a data center design rebuilt layer by layer for real-time AI rather than a retrofit. Model companies serving users buy the same thing enterprises running production AI buy: tokens delivered where demand lives. Five initial production sites are planned; the first came online in July 2026.

What they buy

Grid access first. GridCARE lists AI Fabrik as a customer with 150 MW of capacity under contract, and Raisoni credits the partnership with grid connection in months rather than years. The rest of the buy side: NVIDIA B300 systems by the thousands, modular data center components that assemble faster than traditional builds, and operating experience carried over from Gruve, where the team ran inference infrastructure for enterprise clients before spinning the thesis out into its own company.

Compiled 18 July 2026

Scale proof

Jul 2026
first of five initial production sites comes online; thousands of NVIDIA B300s deploy across the US near tier-1 and tier-2 cities over the following months (aifabrik.com + founder post, Jul 2026, company-provided)
150 MW
power capacity under contract through GridCARE, with grid access for sovereign, low-latency compute arriving in months rather than years, per CEO Tarun Raisoni (gridcare.ai, 2026)
3
founders: Tarun Raisoni (CEO, Gruve), Swati Deshpande, and Tanuj Mohan, an Enlighted co-founder whose smart-building lineage anchors the performance-per-watt focus (Mayfield, Jun 2026)
4
institutional backers from inception: Mayfield, Xora (Temasek), Acclimate Ventures and Cisco Investments, the Gruve investor base following into the spinout. Round size undisclosed (Mayfield + aifabrik.com, Jun–Jul 2026)
The strategy

The bet

Tokens have a distribution problem, and distribution problems get solved at the edge. Training concentrates into a few giant campuses; inference scatters to wherever users, agents and devices sit, and it punishes distance with latency and cost. The CDN precedent drives the whole design. Content delivery moved value from origin servers to distributed edge nodes, and AI Fabrik bets token delivery follows the same path, with modular sites near metros beating central clusters on the metrics that decide production economics: performance per watt and performance per dollar. The Gruve incubation supplies the demand evidence, since the team kept meeting enterprises that had models worth serving and no infrastructure to serve them with. The risks are concrete. One site is live and four are pending, so the operating record spans weeks. Hyperscaler edge zones, CDNs adding GPU capacity, carriers like Telnyx racking GPUs in their own PoPs, and token-delivery software players like Rafay all converge on the same territory. And the moat, fast grid access near cities, doubles as the constraint if interconnection slows.

Deal-making

Natural counterparty

Two buyer types, one product. Model companies serving users at scale need tokens delivered fast in every metro their users occupy. Enterprises running AI in production need the same delivery plus compliance and data locality, which the site-level design targets. On the supply side, AI Fabrik deals with utilities and speed-to-power intermediaries like GridCARE for interconnection, NVIDIA for B300 allocation, and modular construction vendors for the site hardware itself. Gruve sits in a special position as origin, early operating partner and likely demand channel.

Buyer fit

Choose AI Fabrik when

You serve latency-sensitive inference to distributed US users, real-time conversation or real-time decisions, and per-token economics decide whether the product survives. The pitch also fits enterprises with compliance or data-locality needs that central-cloud inference handles badly.

Look elsewhere when

You need training capacity, global coverage, or a long operating record. The network is US-first with five initial sites, the first live only since July 2026, and pricing remains undisclosed. Committing production workloads today means underwriting a build-out still in progress.

Private ledger

Five fields we track but don’t publish

Everything above is public or company-provided, sourced as noted. This card is a draft pending direct provider verification. The fields below change with the build-out and get verified directly with each provider; they live in a private ledger, not on this page.

Site online dates, next 2 quarterslocation · MW · racks · go-live
Effective price floorper-token or per-GPU-hour, at volume
Anchor tenantsnamed or category, by site
Preferred deal structurereserved · on-demand · managed
Capacity-partnerships ownerthe person who signs
Ask about AI Fabrik’s current posture → Answered case-by-case, with the provider’s consent.