NVIDIA B300 nodes.
Blackwell Ultra put 288 GB of HBM3e on one card; a whole 8× B300 node keeps 2.3 TB of model state in memory on one NVLink domain. First nodes ship from October 15, 2026 in Canada; larger blocks from December 2026 in the United States. Bare metal, root access, nothing shared. Written quote within 24 hours.
NVIDIA B300 Blackwell Ultra GPU for HPC
Modern HPC looks less like physics codes and more like trillion-parameter training runs. The Blackwell architecture follows that shift, and the B300 tops it with 288 GB of HBM3e per card. A B300 cluster keeps entire model states in memory, which is what the largest jobs in today's data centers starve for.
Blackwell Ultra changed the math on AI compute, and the B300 puts 288 GB of that math on one card. On the big clouds this card sits behind quotas and capacity reservations; here it is a whole node, allocated to one customer, on a written agreement.
NVIDIA HGX™ B300 node
Eight dedicated cards, single tenant, bare metal.
- 288 GB HBM3e per card — even trillion-parameter models fit on one node
- Fifth-generation NVLink at 1.8 TB/s per GPU
- Root access from the driver up, nothing shared
NVIDIA HGX B300 specifications
The NVIDIA HGX B300 is the reference way to deploy this card: eight B300 GPUs in the SXM6 form factor, wired into one node. The numbers below come from NVIDIA's published materials; the CPU, memory, storage and fabric of your node are named in the written quote.
- GPUs
- 8× NVIDIA B300Blackwell Ultra generation
- GPU memory
- 2.3 TBHBM3e in total, 288 GB per card
- Memory bandwidth
- 8 TB/sper GPU
- Inference
- Up to 144 PFLOPSFP4 sparse · 108 PFLOPS dense FP4
- Training
- 72 PFLOPSFP8 sparse · 36 PFLOPS dense
- Silicon
- 208B transistorsTwo reticle-sized dies, TSMC 4NP
- GPU interconnect
- 1.8 TB/s5th-gen NVLink per GPU
- Networking
- 8× ConnectX-8800 Gb/s each, reference build
- Storage
- ~30 TBlocal NVMe, reference build
Our node and cluster
Every node we reserve is a whole 8-GPU HGX B300 system allocated to one customer: server-class CPUs, DDR5 system memory and local NVMe sized for the GPUs, a dedicated management network and root access to bare metal. No virtualization, no oversubscription, no shared hosts. You run your own stack — drivers, container runtime, orchestration — and we hand over the machine after acceptance benchmarks pass. The CPU model, memory, storage and fabric of your node are confirmed in the written quote and the capacity agreement.
Nodes in one reservation are delivered together as a single cluster on one high-bandwidth fabric, so multi-node training works from day one. Before billing starts every cluster passes the same acceptance run: nccl-tests all_reduce inside the node over NVLink and across nodes over the fabric against the bandwidth thresholds in the agreement, DCGM diagnostics at level 3, and a burn-in. If it fails, the prepayment is refunded.
Where it runs
Tier III facilities in Canada (first nodes, from October 15, 2026) and the United States (larger blocks from December 2026). Whole nodes only, on 1 to 24 month terms. No egress fees. The hardware is of U.S. origin, so U.S. export screening runs during quote acceptance.
Rate
From $6.05/GPU·hr on a 1-year term, $5.30 on 24 months, $7.50 on a single month. Written quote within 24 hours; 30% at signing locks slot and rate; billing starts after acceptance. The full rate card and the published-rate comparison are on /pricing.
Best use cases for NVIDIA B300
The B300 is a specialist card, and it pays off fastest where memory decides the outcome. Below is the map, from training on large datasets to the inference patterns that keep a 288 GB card busy.
LLM training & inference
288 GB per card and 8 TB/s bandwidth make the B300 the default choice across frontier training and inference.
- Dense models past 500B parameters with fewer pipeline-parallel stages
- Mixture-of-experts architectures with expert weights resident in HBM3e
- Trillion-parameter models served from a single eight-GPU node
- High-concurrency APIs where memory, not FLOPS, sets the batch ceiling
Fine-tuning
With 288 GB per card, full fine-tuning of 70B-class models runs comfortably inside a single node.
- Full-parameter fine-tuning of 70B+ models inside one node
- LoRA and QLoRA runs batched across several teams
- Domain adaptation on proprietary corpora that can't leave a dedicated environment
- RLHF and DPO pipelines with the reward model in memory next to the policy
AI agents
Agents carry state for hours, which makes memory capacity the real budget line.
- Multi-agent systems sharing one resident base model
- Long-horizon tasks that accumulate context over hours of tool calls
- Reasoning-heavy agents burning thinking tokens at every step
- Around-the-clock workloads on hardware nobody else touches
Scientific computing
Science on the B300 means AI-shaped science: foundation models trained on large experimental datasets.
- Protein structure and drug discovery models in the AlphaFold lineage
- Weather and climate emulators trained on decades of observations
- Materials discovery with graph networks over simulation archives
- Genomics models that read sequences too long for smaller cards
RAG
A production RAG stack runs several models side by side, and the B300 hosts all of them without eviction.
- Embedding and reranking models colocated with the generator
- GPU-accelerated vector search over million-document indexes
- Long-context synthesis across dozens of retrieved passages
- Enterprise assistants grounded in private corpora on dedicated hardware
Video & image generation
Generative media is a memory story: frames are large and video models are larger.
- Diffusion transformers for high-resolution, multi-second video
- Image models served at production pace for user-facing products
- Batch rendering queues for creative and marketing pipelines
- Custom style models trained on studio archives
How the B300 compares to other popular GPUs
Memory and bandwidth decide the fit. Here's the B300 next to the cards teams cross-shop.
| Spec | NVIDIA B300 | NVIDIA B200 | NVIDIA H200 | NVIDIA H100 (SXM) |
|---|---|---|---|---|
| GPU memory | 288 GB HBM3e | 180 GB HBM3e | 141 GB HBM3e | 80 GB HBM3 |
| Memory bandwidth | ~8 TB/s | ~8 TB/s | ~4.8 TB/s | ~3.35 TB/s |
| NVLink per GPU | 1.8 TB/s (gen 5) | 1.8 TB/s (gen 5) | 900 GB/s (gen 4) | 900 GB/s (gen 4) |
| Lowest supported precision | FP4 | FP4 | FP8 | FP8 |
| FP4 inference, 8-GPU node | 144 PFLOPS sparse · 108 dense | 144 PFLOPS sparse · 72 dense | — | — |
| Full fine-tuning of 70B-class models on one node | Yes | Yes | No | No |
Why reserve instead of buying
Buying a Blackwell Ultra node ties up capital for years; reserving the same hardware turns it into a rate per GPU-hour on a 1 to 24 month term.
Power, cooling and failed-card replacements sit inside the rate, so your operational costs stay flat and predictable. No egress fees.
One month to two years, written into the capacity agreement. The rate holds for the whole term; nothing is metered in the background.
When the next NVIDIA generation ships, you move to it under a new agreement while owners start a new procurement cycle.
Before you reserve
Is the reservation paid or binding?
Requesting a configuration is free and creates no obligation. Your quote — site, delivery window, spec and rate — comes within 24 hours. A 30% prepayment then locks your slot and rate; from that point the reservation is take-or-pay.
Can I cancel a locked reservation?
A locked reservation is a rental contract, not a hold: the 30% prepayment is credited to your term and is not refunded if you cancel. Until you accept a quote, nothing is owed.
What if the delivery window slips?
If the delivery window in your agreement is missed, you choose: a rate reduction for the delay, or a full refund of the prepayment if the delay is material. Billing never starts before acceptance, and your rate does not move.
What if the cluster doesn't pass acceptance?
Acceptance criteria are written into your agreement before delivery. If the cluster doesn't pass them, the prepayment is refunded in full. The criteria are spec-based until first delivery: nccl-tests all_reduce inside the node over NVLink and across nodes over the cluster fabric, against the bandwidth thresholds named in the agreement; DCGM diagnostics at level 3; and a burn-in run.
Ready to reserve?
Written quote in 24h. No payment until you accept.
Questions before a request: cloud@theai.com