
We are the deployment arm for heterogeneous compute.
Backed by
The hardware
This is what you actually buy from us: dedicated racks of purpose-built silicon, one contract, sited with your fleet. Specs are the vendor's own; where a number isn't published, we say so.
Drag to rotate · 42U
Priced per contract, sized to term, volume and site.
Speed
A chat turn is one pass through a model. An agent run is hundreds or thousands, in sequence, each waiting on the one before it. Latency stops averaging out and starts compounding into wall-clock time.
100–1,000×
model passes per agent task, versus one for a chat turn.
OpenCode running the identical task on both sides. A real coding session at measured operating points.

Prefill vs decode
Prefill is compute-bound and parallel — exactly what a GPU is built for. Decode is memory-bound and autoregressive: one token at a time, limited by bandwidth, not FLOPs. Adding compute does not move it, and the ceiling drops as the model gets larger.
The B300 ceiling sits under
~320tokens / s / user
and materially lower on frontier-scale models. Above it, the GPU cannot operate.
Interactivity frontier — MiniMax M2.7
Output throughput per rack against the per-user token rate it sustains.
Everything to the right of the line is a market only purpose-built decode silicon can serve.
GC internal benchmark set, SN50 20kW / 16-RDU decode rack, B300 approximated from SemiAnalysis InferenceX.
Disaggregation
Prefill
NVIDIA B300 tier
compute-bound, parallel
Decode
SambaNova SN50 · Cerebras · Positron · d‑Matrix
memory-bound, autoregressive
Your surface
Bare metal, or one endpoint if you want it
one contract, one set of SLAs
The offer
We own the racks, we do the bring-up, and you get dedicated silicon under contract, sited with your fleet, with root on the machines you're paying for.
You pick a latency target and a price. We put the hardware behind it.
Prefill
NVIDIA B300
Compute-bound. We bring the allocation, so you keep the stack you know.
Fast
SambaNova SN50 · Positron · d‑Matrix
1–2k tok/s/user on multi-trillion-parameter models. SN50 in production first; Positron and d‑Matrix widen the ladder as they reach volume.
Fastest
Cerebras
SRAM wafer-scale, at the top of the ladder for latency-critical work.
What it takes
Siting, bring-up and operations are on us. You are buying capacity under contract, not standing up a build.
In our colocation pipeline
30MW+
We buy the racks and carry them on our balance sheet. You get dedicated silicon under contract, with root on the machines.
Price-protected quota across three vendors, placed ahead of merchant availability.
We handle colocation site selection, power and cooling in existing US facilities, matched to each vendor's requirements. Capacity sits alongside your fleet where possible.
Model bring-up on closed vendor stacks, and orchestration of prefill and decode as one service, for the life of the contract.
Switching
Every switching objection assumes we are asking you to abandon a stack you have already paid to build. We are not. The prefill half stays on the GPUs you want, we bring the allocation, and the decode half moves to silicon that can actually serve it.
“A new chip means a new software stack.”
We do the model bring-up. One OpenAI-compatible API spans prefill and decode; bare metal if you want it.
“We need day-zero model support.”
Written into the SLA as a contracted bring-up commitment with a defined window.
“Our workloads are lumpy. Utilisation risk is ours.”
Reserved decode capacity with burst onto the paired B300 fleet. You are never decode-only, so a quiet week is not a stranded rack.
“We'd be locked to one vendor's chip.”
We hold quota across every vendor, so you are not tied to one chip. Today you are tied to one.
The business
Tell us the model, the traffic shape and the interactivity target. We'll come back with a rate card and a deployment slot — not a waitlist.