
Viacheslav Zhvart.
CTO at TheAI. Building the B300 cloud this blog is about.
Articles.
How Much Blackwell Actually Exists?
Blackwell shipments set records every quarter — and the median public B300 on-demand price rose 57% since November. We reconstruct how much has actually shipped (~7M packages), how little of it is rentable (two sub-12-month listings), and why prices rise while supply booms.
The Blackwell Fact Table
Every figure from “How Much Blackwell Actually Exists?” with its source: shipments, the B300 slice, capex and anchor deals, manufacturing ceilings, Rubin, the H100 precedent — plus methodology notes and unit conventions.
Kimi K3 (2.8T) and six other open models on a single 8×B300 node — full vLLM numbers
We ran seven current open models — from Gemma 4 31B to Kimi K3's 2.78T parameters — through out-of-box vLLM on one 8× B300 node. TTFT, throughput, what didn't launch, and what surprised us.
We asked ~50 GPU providers for B300 pricing. Here's the map
Before building our own cloud, we sent the same question to about fifty GPU providers: 8× B300 nodes, short-to-mid term — what's your price? This is what the market actually looks like.
Whole-node vs fractional GPU rental — how to choose
Fractional GPUs, MIG, vGPU and shared hosts look cheaper by the hour. Here is when they cost you more, and when a whole bare-metal node is the right unit.
Multi-node training — what the fabric actually does
When training outgrows one node, the network sets your scaling efficiency. A practical look at NVLink, InfiniBand, collective operations and non-blocking topology.
The economics of large-model inference on GPUs
Why inference cost is set by memory bandwidth and the KV cache, not raw FLOPs — and how batching, quantization and VRAM headroom decide your cost per token.