Blog.
GPU economics, inference, multi-node training and bare-metal infrastructure — notes from the team running whole-node B300.
We asked ~50 GPU providers for B300 pricing. Here's the map
Before building our own cloud, we sent the same question to about fifty GPU providers: 8× B300 nodes, short-to-mid term — what's your price? This is what the market actually looks like.
Whole-node vs fractional GPU rental — how to choose
Fractional GPUs, MIG, vGPU and shared hosts look cheaper by the hour. Here is when they cost you more, and when a whole bare-metal node is the right unit.
The economics of large-model inference on GPUs
Why inference cost is set by memory bandwidth and the KV cache, not raw FLOPs — and how batching, quantization and VRAM headroom decide your cost per token.