NVIDIA B200 nodes.
The workhorse Blackwell part: 180 GB of HBM3e per card, 1,440 GB per node, the same NVLink 5 domain as B300 at a lower rate per GPU-hour. Delivery December 2026 in the United States, a liquid-cooled Tier III facility, up to 32 nodes in one NDR InfiniBand fabric. Bare metal, root access, nothing shared. Written quote within 24 hours.
NVIDIA B200 GPUs for high-performance computing
The B200 packs two dies and 208 billion transistors into a single GPU built on the Blackwell architecture. Fifth-generation Tensor Cores add native FP4 support, roughly doubling inference throughput on quantized models compared with FP8.
The memory system matters just as much: at about 8 TB/s, the card keeps large models fed where older generations stalled on reads. Most serious budgets in the Blackwell generation land on the B200: near-flagship throughput without the scarcity, at a lower rate per GPU-hour than the B300.
NVIDIA HGX™ B200 node
Eight dedicated cards, single tenant, bare metal.
- 180 GB HBM3e per card — most open-weight models fit on one node
- Fifth-generation NVLink at 1.8 TB/s per GPU
- Root access from the driver up, nothing shared
NVIDIA DGX B200 specifications
The DGX B200 is NVIDIA's reference eight-GPU system and a useful yardstick for what a B200 node should deliver. Our nodes are HGX B200 systems from the batch described below; the numbers here come from NVIDIA's published materials.
- GPUs
- 8× NVIDIA B200Tensor Core GPUs
- GPU memory
- 1,440 GBHBM3e per node · 180 GB per card
- Memory bandwidth
- 8 TB/sper GPU
- AI performance
- 72 / 144 PFLOPSFP8 training · FP4 inference
- GPU interconnect
- 1.8 TB/s5th-gen NVLink per GPU
- CPUs
- 2× Xeon 8570112 cores total, reference build
- System memory
- up to 4 TBDDR5, reference build
- Networking
- 8× ConnectX-7400 Gb/s each, reference build
- Storage
- ~30 TBlocal NVMe, reference build
Our node and cluster
Our B200 nodes are HGX B200 systems from one batch: eight B200 180GB HBM3e GPUs (1.44 TB per node), two high-performance server-class CPUs, about 2 TB of DDR5, roughly 31 TB of local enterprise NVMe plus dedicated boot storage, a dedicated out-of-band management network, and root access to bare metal. Direct liquid cooling at rack level with a dedicated CDU.
The CPU model is confirmed in the written quote; the full configuration is fixed in the capacity agreement.
Nodes connect over NDR InfiniBand on a dedicated leaf/spine fabric, and one reservation can span up to 32 nodes /256 GPUs delivered together as a single cluster. Before billing starts every cluster passes the same acceptance run: nccl-tests all_reduce inside the node over NVLink and across nodes over InfiniBand against the bandwidth thresholds in the agreement, DCGM diagnostics at level 3, and a burn-in. If it fails, the prepayment is refunded.
Where it runs
A liquid-cooled Tier III facility in the United States, delivery December 2026. Whole nodes only, on 1 to 24 month terms; whole-lot 36-month pricing on request. No egress fees. The hardware is of U.S. origin, so U.S. export screening runs during quote acceptance.
Rate
From $4.85/GPU·hr on a 1-year term, $4.50 on 24 months, $5.75 on a single month. Written quote within 24 hours; 30% at signing locks slot and rate; billing starts after acceptance. The full rate card and the published-rate comparison are on /pricing.
Best use cases for NVIDIA B200
Memory decides what fits on a card, and bandwidth decides how fast it runs. The B200 covers both, which is why the list below stretches from training to generative serving.
LLM training & inference
FP4 throughput and fast memory make the B200 a strong default for the whole model lifecycle.
- Foundation models with 100B+ parameters
- Mixture-of-experts (MoE) architectures
- Long-context serving with heavy KV caches
- Batch inference at FP4 precision
Fine-tuning
With 180 GB per card, most open-weight models fine-tune on a single node.
- Full-parameter fine-tuning of 70B-class models
- LoRA and QLoRA adaptation
- Domain adaptation on proprietary corpora
- RLHF and DPO alignment passes
AI agents
Agent fleets run around the clock and need inference capacity that holds under bursty tool calls.
- Multi-agent systems on one backend
- Code agents with long context
- Customer-facing assistants under live traffic
- Agent evaluation and simulation
Scientific computing
Simulation workloads that pair AI surrogate models with classic numerical methods map well onto Blackwell.
- Scientific foundation models for drug discovery
- Molecular dynamics with ML force fields
- Climate and weather model emulation
- Genomics pipelines with deep learning stages
RAG
Retrieval-augmented generation lives and dies on latency; fast memory keeps embedder and generator responsive.
- Embedding generation at corpus scale
- Reranking models in the retrieval path
- Long retrieved contexts, low latency
- Thousands of concurrent queries
Video & image generation
Diffusion and video models are memory gluttons; 180 GB per card means fewer compromises.
- Text-to-video at production resolution
- Diffusion pipelines with large batches
- Image upscaling and restoration
- Frame interpolation and video editing
How the B200 compares to other popular GPUs
Memory and bandwidth decide the fit. Here's the B200 next to the cards teams cross-shop.
| Spec | NVIDIA B200 | NVIDIA B300 | NVIDIA H200 | NVIDIA H100 (SXM) |
|---|---|---|---|---|
| Architecture | Blackwell | Blackwell Ultra | Hopper | Hopper |
| GPU memory | 180 GB HBM3e | 288 GB HBM3e | 141 GB HBM3e | 80 GB HBM3 |
| Memory bandwidth | ~8 TB/s | ~8 TB/s | ~4.8 TB/s | ~3.35 TB/s |
| NVLink per GPU | 1.8 TB/s (gen 5) | 1.8 TB/s (gen 5) | 900 GB/s (gen 4) | 900 GB/s (gen 4) |
| Native FP4 | Yes | Yes | No | No |
| Fits best | Training + high-throughput inference | Memory-bound and long-context serving | Memory-bound inference | Established Hopper workloads |
Why reserve instead of buying
A dedicated cluster lands as a rate per GPU-hour on a 1 to 24 month term; the capex conversation never happens.
Start at one node and grow to a multi-node block on the same fabric, inside the same agreement.
A fixed rate turns cost per token into a spreadsheet exercise. No usage meters underneath, no egress fees.
Bare metal and root access, your own stack from driver to container runtime. Nothing changes underneath you mid-training.
Before you reserve
Is the reservation paid or binding?
Requesting a configuration is free and creates no obligation. Your quote — site, delivery window, spec and rate — comes within 24 hours. A 30% prepayment then locks your slot and rate; from that point the reservation is take-or-pay.
Can I cancel a locked reservation?
A locked reservation is a rental contract, not a hold: the 30% prepayment is credited to your term and is not refunded if you cancel. Until you accept a quote, nothing is owed.
What if the delivery window slips?
If the delivery window in your agreement is missed, you choose: a rate reduction for the delay, or a full refund of the prepayment if the delay is material. Billing never starts before acceptance, and your rate does not move.
What if the cluster doesn't pass acceptance?
Acceptance criteria are written into your agreement before delivery. If the cluster doesn't pass them, the prepayment is refunded in full. The criteria are spec-based until first delivery: nccl-tests all_reduce inside the node over NVLink and across nodes over the cluster fabric, against the bandwidth thresholds named in the agreement; DCGM diagnostics at level 3; and a burn-in run.
Ready to reserve?
Written quote in 24h. No payment until you accept.
Questions before a request: cloud@theai.com