TheAI·cloud
CA-CENTRAL-1 ● OPENING Q4 2026FIRST RACK — 10 NODESRESERVING NOW

Blog.

GPU economics, inference, multi-node training and bare-metal infrastructure — notes from the team running whole-node B300.

BLACKWELL · B300 · GPU-MARKET

How Much Blackwell Actually Exists?

Blackwell shipments set records every quarter — and the median public B300 on-demand price rose 57% since November. We reconstruct how much has actually shipped (~7M packages), how little of it is rentable (two sub-12-month listings), and why prices rise while supply booms.

Aug 27, 2026 · 13 min read
BLACKWELL · B300 · SOURCES

The Blackwell Fact Table

Every figure from “How Much Blackwell Actually Exists?” with its source: shipments, the B300 slice, capex and anchor deals, manufacturing ceilings, Rubin, the H100 precedent — plus methodology notes and unit conventions.

Aug 27, 2026 · 17 min read
BENCHMARKS · B300 · VLLM

Kimi K3 (2.8T) and six other open models on a single 8×B300 node — full vLLM numbers

We ran seven current open models — from Gemma 4 31B to Kimi K3's 2.78T parameters — through out-of-box vLLM on one 8× B300 node. TTFT, throughput, what didn't launch, and what surprised us.

Aug 12, 2026 · 8 min read
GPU-ECONOMICS · GPU-RENTAL · B300

We asked ~50 GPU providers for B300 pricing. Here's the map

Before building our own cloud, we sent the same question to about fifty GPU providers: 8× B300 nodes, short-to-mid term — what's your price? This is what the market actually looks like.

Aug 10, 2026 · 6 min read
BARE-METAL · GPU-RENTAL · GPU-ECONOMICS

Whole-node vs fractional GPU rental — how to choose

Fractional GPUs, MIG, vGPU and shared hosts look cheaper by the hour. Here is when they cost you more, and when a whole bare-metal node is the right unit.

Aug 7, 2026 · 5 min read
MULTI-NODE · INFINIBAND · BARE-METAL

Multi-node training — what the fabric actually does

When training outgrows one node, the network sets your scaling efficiency. A practical look at NVLink, InfiniBand, collective operations and non-blocking topology.

Aug 7, 2026 · 5 min read
INFERENCE · GPU-ECONOMICS · LLM

The economics of large-model inference on GPUs

Why inference cost is set by memory bandwidth and the KV cache, not raw FLOPs — and how batching, quantization and VRAM headroom decide your cost per token.

Aug 7, 2026 · 5 min read
Blog — GPU cloud, inference & infrastructure | TheAI Cloud