We secured a $400M debt facility with Upper90 to scale inference compute.Read

We deploySambaNova

We are the deployment arm for heterogeneous compute.

Backed by

Village GlobalFuseCaryaUpper90

The hardware

Pick a rack. See what's in it.

This is what you actually buy from us: dedicated racks of purpose-built silicon, one contract, sited with your fleet. Specs are the vendor's own; where a number isn't published, we say so.

Shipping H2 2026

Drag to rotate · 42U

Decode rack

SambaNova SambaRack SN50

Sixteen reconfigurable dataflow units in a standard air-cooled footprint. Our first decode rack in production.

Accelerators
16 RDUs · 2 nodes × 8
Power
20 kW average per rack
Cooling
Air-cooled — drops into existing colo
Memory
64 GB HBM2E + up to 2 TB DDR5 per RDU
Speed
895 tok/s per user · Llama 3.3 70B
Model ceiling
Up to 10T parameters, 10M-token context
Interconnect
800GbE scale-up · 2.2 TB/s chip-to-chip
List price
Not public
Request a quote

Priced per contract, sized to term, volume and site.

Speed

Agents make decode speed the product.

A chat turn is one pass through a model. An agent run is hundreds or thousands, in sequence, each waiting on the one before it. Latency stops averaging out and starts compounding into wall-clock time.

100–1,000×

model passes per agent task, versus one for a chat turn.

Coding agents

Hours of work compressed into minutes. Serial tool calls make decode speed the binding constraint.

Up to 16.1x faster

OpenCode running the identical task on both sides. A real coding session at measured operating points.

OpenCodeGPU cloud today
110 tok/s/user
loop 1 of 32 · 0 tokensstreaming…
OpenCodeGeneral Compute
2,000 tok/s/user
loop 1 of 32 · 0 tokensstreaming…

Prefill vs decode

GPUs are prefill machines. Decode is where they stall.

Prefill is compute-bound and parallel — exactly what a GPU is built for. Decode is memory-bound and autoregressive: one token at a time, limited by bandwidth, not FLOPs. Adding compute does not move it, and the ceiling drops as the model gets larger.

The B300 ceiling sits under

~320tokens / s / user

and materially lower on frontier-scale models. Above it, the GPU cannot operate.

Interactivity frontier — MiniMax M2.7

Output throughput per rack against the per-user token rate it sustains.

SambaNova SN50NVIDIA B300

Everything to the right of the line is a market only purpose-built decode silicon can serve.

GC internal benchmark set, SN50 20kW / 16-RDU decode rack, B300 approximated from SemiAnalysis InferenceX.

Disaggregation

We split the workload. Prefill on GPUs, decode on ASICs.

Prefill

NVIDIA B300 tier

compute-bound, parallel

Decode

SambaNova SN50 · Cerebras · Positron · d‑Matrix

memory-bound, autoregressive

Your surface

Bare metal, or one endpoint if you want it

one contract, one set of SLAs

The offer

We buy the chips. You get the metal.

We own the racks, we do the bring-up, and you get dedicated silicon under contract, sited with your fleet, with root on the machines you're paying for.

You pick a latency target and a price. We put the hardware behind it.

Prefill

NVIDIA B300

Compute-bound. We bring the allocation, so you keep the stack you know.

Fast

SambaNova SN50 · Positron · d‑Matrix

1–2k tok/s/user on multi-trillion-parameter models. SN50 in production first; Positron and d‑Matrix widen the ladder as they reach volume.

Fastest

Cerebras

SRAM wafer-scale, at the top of the ladder for latency-critical work.

What it takes

Hardware, allocation, data centers, software. We run all four.

Siting, bring-up and operations are on us. You are buying capacity under contract, not standing up a build.

In our colocation pipeline

30MW+

01

Hardware

We buy the racks and carry them on our balance sheet. You get dedicated silicon under contract, with root on the machines.

02

Allocation

Price-protected quota across three vendors, placed ahead of merchant availability.

03

Data centers

We handle colocation site selection, power and cooling in existing US facilities, matched to each vendor's requirements. Capacity sits alongside your fleet where possible.

04

Software

Model bring-up on closed vendor stacks, and orchestration of prefill and decode as one service, for the life of the contract.

Switching

We don't take GPUs away. We bring the allocation for prefill.

Every switching objection assumes we are asking you to abandon a stack you have already paid to build. We are not. The prefill half stays on the GPUs you want, we bring the allocation, and the decode half moves to silicon that can actually serve it.

A new chip means a new software stack.

We do the model bring-up. One OpenAI-compatible API spans prefill and decode; bare metal if you want it.

We need day-zero model support.

Written into the SLA as a contracted bring-up commitment with a defined window.

Our workloads are lumpy. Utilisation risk is ours.

Reserved decode capacity with burst onto the paired B300 fleet. You are never decode-only, so a quiet week is not a stranded rack.

We'd be locked to one vendor's chip.

We hold quota across every vendor, so you are not tied to one chip. Today you are tied to one.

The business

Every chip that is not a GPU needs a route to market. We're building it.

Tell us the model, the traffic shape and the interactivity target. We'll come back with a rate card and a deployment slot — not a waitlist.

ModeHumanAgent