Rung 6 · Tokens, custom models

FriendliAI

Sells tokens from your model, and buys the GPUs to do it.

Public facts compiled 4 July 2026 · awaiting FriendliAI's review  ·  official site: friendli.ai

HQ
Redwood City / SF, CA
Center of gravity
US · Korea / APAC
Silicon
NVIDIA via partners (B300)
Contract shape
Per-token · dedicated · BYO-cluster
The two-sided view

What FriendliAI sells, and what it buys

Atomic unit sold

Tokens and inference infrastructure in three shapes: serverless per-token endpoints over open models, Dedicated Endpoints for custom and fine-tuned models, and Friendli Container, the same engine running inside a customer’s own GPU cluster. 540,000+ Hugging Face models deploy one-click; 99.99% uptime SLA on geo-distributed infrastructure.

What they buy

GPU capacity across regions and clouds. The design stays asset-light: FriendliAI sources the fleets behind its SLA rather than owning them, and the 2026 Samsung Cloud Platform alliance for NVIDIA B300 inference capacity is the visible tip. This is the clearest demand-side posture of any card on this ladder: FriendliAI is a perpetual GPU buyer.

Compiled 4 July 2026

Scale proof

6–7×
expected 2025 revenue over 2024, on 25–30 large enterprise clients (CEO Byung-Gon Chun, Aug 2025), with strong gross margins while scaling
~$26M
total raised: a $20M seed extension (Aug 2025) led by Capstone Partners with Sierra Ventures, KDB and KB Securities. Inference peers raise ten times more
540K+
Hugging Face models deployable one-click; exclusive API provider of LG AI Research’s EXAONE; customers include Scatter Lab (Zeta) and Upstage (Solar)
2026
Samsung Cloud Platform strategic alliance for NVIDIA B300 inference; former Moloco COO Brian Yoo joins as Chief Business Officer; San Francisco expansion
The strategy

The bet

Inference is a software problem. That's the wager, and real estate is the counter-position. Founder Byung-Gon Chun invented continuous batching (published as Orca at OSDI 2022, now industry standard in LLM serving), and the company bets that custom kernels, caching and speculative decoding squeeze enough extra tokens out of each GPU that Friendli can stay asset-light and rent capacity while competitors buy data centers. DeepInfra shares the rung and takes the opposite posture, which gives the tokens rung two ways to sell the same unit.

Deal-making

Natural counterparty

On the sell side: enterprises with custom or fine-tuned models (the LG EXAONE pattern) and APAC-rooted AI companies needing regional serving. On the buy side: capacity providers. Because Friendli procures GPUs across regions and price points continuously, it is a natural demand-side counterparty for every supply-side company on this ladder. The APAC overlap makes GMI Cloud the most obvious pairing.

Buyer fit

Choose FriendliAI when

You serve a custom or fine-tuned model and want frontier-grade latency and cost without owning infrastructure, or you want the inference engine deployed inside your own GPU cluster.

Look elsewhere when

You just need cheap catalog-model tokens at massive scale and don’t care whose engine serves them.

Private ledger

Five fields we track but don’t publish

Everything above is public and compiled from primary sources. The fields below are demand-side: what Friendli is looking for, verified in conversation and never published.

Capacity sought, next 90 daysGPU type · region · volume · price ceiling
Effective price ceilingwhat they'll pay per GPU-hour, at volume
Current serving regionswhere the fleet actually runs
Preferred deal structurecapacity terms · co-sell terms
Capacity-procurement ownerthe person who signs
Ask about FriendliAI’s current posture → Answered case-by-case, with the company’s consent.