TheAI·cloud
NVIDIA BLACKWELL ULTRA · B300 · WHOLE 8-GPU NODES

NVIDIA B300 nodes.

Blackwell Ultra put 288 GB of HBM3e on one card; a whole 8× B300 node keeps 2.3 TB of model state in memory on one NVLink domain. First nodes ship from October 15, 2026 in Canada; larger blocks from December 2026 in the United States. Bare metal, root access, nothing shared. Written quote within 24 hours.

288 GBHBM3e per card2,304 GB across the 8-card node
8 TB/sMemory bandwidthKeeps large models fed, no read stalls
Native FP45th-gen Tensor Cores~2× inference throughput vs FP8
1.8 TB/sNVLink per GPUFifth-generation NVLink, all-to-all inside the node

NVIDIA B300 Blackwell Ultra GPU for HPC

Modern HPC looks less like physics codes and more like trillion-parameter training runs. The Blackwell architecture follows that shift, and the B300 tops it with 288 GB of HBM3e per card. A B300 cluster keeps entire model states in memory, which is what the largest jobs in today's data centers starve for.

Blackwell Ultra changed the math on AI compute, and the B300 puts 288 GB of that math on one card. On the big clouds this card sits behind quotas and capacity reservations; here it is a whole node, allocated to one customer, on a written agreement.

208B transistorsTwo dies on Blackwell Ultra
FP4 Tensor Cores~2× inference vs FP8
~8 TB/s memoryNo read stalls at scale

NVIDIA HGX™ B300 node

Eight dedicated cards, single tenant, bare metal.

NVIDIA HGX B300 specifications

The NVIDIA HGX B300 is the reference way to deploy this card: eight B300 GPUs in the SXM6 form factor, wired into one node. The numbers below come from NVIDIA's published materials; the CPU, memory, storage and fabric of your node are named in the written quote.

GPUs
8× NVIDIA B300Blackwell Ultra generation
GPU memory
2.3 TBHBM3e in total, 288 GB per card
Memory bandwidth
8 TB/sper GPU
Inference
Up to 144 PFLOPSFP4 sparse · 108 PFLOPS dense FP4
Training
72 PFLOPSFP8 sparse · 36 PFLOPS dense
Silicon
208B transistorsTwo reticle-sized dies, TSMC 4NP
GPU interconnect
1.8 TB/s5th-gen NVLink per GPU
Networking
8× ConnectX-8800 Gb/s each, reference build
Storage
~30 TBlocal NVMe, reference build

Our node and cluster

Every node we reserve is a whole 8-GPU HGX B300 system allocated to one customer: server-class CPUs, DDR5 system memory and local NVMe sized for the GPUs, a dedicated management network and root access to bare metal. No virtualization, no oversubscription, no shared hosts. You run your own stack — drivers, container runtime, orchestration — and we hand over the machine after acceptance benchmarks pass. The CPU model, memory, storage and fabric of your node are confirmed in the written quote and the capacity agreement.

Nodes in one reservation are delivered together as a single cluster on one high-bandwidth fabric, so multi-node training works from day one. Before billing starts every cluster passes the same acceptance run: nccl-tests all_reduce inside the node over NVLink and across nodes over the fabric against the bandwidth thresholds in the agreement, DCGM diagnostics at level 3, and a burn-in. If it fails, the prepayment is refunded.

Where it runs

Tier III facilities in Canada (first nodes, from October 15, 2026) and the United States (larger blocks from December 2026). Whole nodes only, on 1 to 24 month terms. No egress fees. The hardware is of U.S. origin, so U.S. export screening runs during quote acceptance.

Rate

From $6.05/GPU·hr on a 1-year term, $5.30 on 24 months, $7.50 on a single month. Written quote within 24 hours; 30% at signing locks slot and rate; billing starts after acceptance. The full rate card and the published-rate comparison are on /pricing.

Best use cases for NVIDIA B300

The B300 is a specialist card, and it pays off fastest where memory decides the outcome. Below is the map, from training on large datasets to the inference patterns that keep a 288 GB card busy.

LLM training & inference

288 GB per card and 8 TB/s bandwidth make the B300 the default choice across frontier training and inference.

  • Dense models past 500B parameters with fewer pipeline-parallel stages
  • Mixture-of-experts architectures with expert weights resident in HBM3e
  • Trillion-parameter models served from a single eight-GPU node
  • High-concurrency APIs where memory, not FLOPS, sets the batch ceiling

Fine-tuning

With 288 GB per card, full fine-tuning of 70B-class models runs comfortably inside a single node.

  • Full-parameter fine-tuning of 70B+ models inside one node
  • LoRA and QLoRA runs batched across several teams
  • Domain adaptation on proprietary corpora that can't leave a dedicated environment
  • RLHF and DPO pipelines with the reward model in memory next to the policy

AI agents

Agents carry state for hours, which makes memory capacity the real budget line.

  • Multi-agent systems sharing one resident base model
  • Long-horizon tasks that accumulate context over hours of tool calls
  • Reasoning-heavy agents burning thinking tokens at every step
  • Around-the-clock workloads on hardware nobody else touches

Scientific computing

Science on the B300 means AI-shaped science: foundation models trained on large experimental datasets.

  • Protein structure and drug discovery models in the AlphaFold lineage
  • Weather and climate emulators trained on decades of observations
  • Materials discovery with graph networks over simulation archives
  • Genomics models that read sequences too long for smaller cards

RAG

A production RAG stack runs several models side by side, and the B300 hosts all of them without eviction.

  • Embedding and reranking models colocated with the generator
  • GPU-accelerated vector search over million-document indexes
  • Long-context synthesis across dozens of retrieved passages
  • Enterprise assistants grounded in private corpora on dedicated hardware

Video & image generation

Generative media is a memory story: frames are large and video models are larger.

  • Diffusion transformers for high-resolution, multi-second video
  • Image models served at production pace for user-facing products
  • Batch rendering queues for creative and marketing pipelines
  • Custom style models trained on studio archives

How the B300 compares to other popular GPUs

Memory and bandwidth decide the fit. Here's the B300 next to the cards teams cross-shop.

SpecNVIDIA B300NVIDIA B200NVIDIA H200NVIDIA H100 (SXM)
GPU memory288 GB HBM3e180 GB HBM3e141 GB HBM3e80 GB HBM3
Memory bandwidth~8 TB/s~8 TB/s~4.8 TB/s~3.35 TB/s
NVLink per GPU1.8 TB/s (gen 5)1.8 TB/s (gen 5)900 GB/s (gen 4)900 GB/s (gen 4)
Lowest supported precisionFP4FP4FP8FP8
FP4 inference, 8-GPU node144 PFLOPS sparse · 108 dense144 PFLOPS sparse · 72 dense——
Full fine-tuning of 70B-class models on one nodeYesYesNoNo

Why reserve instead of buying

No capital outlay

Buying a Blackwell Ultra node ties up capital for years; reserving the same hardware turns it into a rate per GPU-hour on a 1 to 24 month term.

Flat operating costs

Power, cooling and failed-card replacements sit inside the rate, so your operational costs stay flat and predictable. No egress fees.

Terms that fit the roadmap

One month to two years, written into the capacity agreement. The rate holds for the whole term; nothing is metered in the background.

Always current hardware

When the next NVIDIA generation ships, you move to it under a new agreement while owners start a new procurement cycle.

Before you reserve

Is the reservation paid or binding?

Requesting a configuration is free and creates no obligation. Your quote — site, delivery window, spec and rate — comes within 24 hours. A 30% prepayment then locks your slot and rate; from that point the reservation is take-or-pay.

Can I cancel a locked reservation?

A locked reservation is a rental contract, not a hold: the 30% prepayment is credited to your term and is not refunded if you cancel. Until you accept a quote, nothing is owed.

What if the delivery window slips?

If the delivery window in your agreement is missed, you choose: a rate reduction for the delay, or a full refund of the prepayment if the delay is material. Billing never starts before acceptance, and your rate does not move.

What if the cluster doesn't pass acceptance?

Acceptance criteria are written into your agreement before delivery. If the cluster doesn't pass them, the prepayment is refunded in full. The criteria are spec-based until first delivery: nccl-tests all_reduce inside the node over NVLink and across nodes over the cluster fabric, against the bandwidth thresholds named in the agreement; DCGM diagnostics at level 3; and a burn-in run.

All questions →

Ready to reserve?

Written quote in 24h. No payment until you accept.

Questions before a request: cloud@theai.com

NVIDIA B300 nodes — Blackwell Ultra, whole 8-GPU nodes | TheAI Cloud