TheAI·cloud
CA-CENTRAL-1 ● OPENING Q4 2026FIRST RACK — 10 NODESRESERVING NOW

How Much Blackwell Actually Exists?

Blackwell shipments set records every quarter — and the median public B300 on-demand price rose 57% since November. We reconstruct how much has actually shipped (~7M packages), how little of it is rentable (two sub-12-month listings), and why prices rise while supply booms.

· 13 min read

Everyone prices B300 capacity. Almost nobody asks how much of it physically exists — and the two numbers that matter (how much has shipped, and how much you can actually rent) turn out to be wildly different.

Here's the paradox in one line: Blackwell shipments set records every quarter, hyperscaler capex nearly doubled — and the median public on-demand price for a B300 rose 57% since November, from $5.00 to $7.87/GPU·hr (getdeploying, Aug 18 snapshot). And no, it's not an index artifact: the rise is the same offers repricing, not expensive new listings joining — decomposition below. Booming supply, rising prices. This post reconstructs why, with sources for every number — the full fact table is linked at the end.

First, the unit problem

NVIDIA reports revenue, not units — and when units do get mentioned, the unit itself is slippery. A Blackwell package contains two silicon dies, and NVIDIA has started counting dies as "GPUs" (Huang, GTC March 2025: "each GPU die is a GPU"). So when Jensen Huang says "6 million Blackwells shipped" (GTC DC, October 2025), that's dies — roughly 3.0 million packages, the thing you'd actually recognize as a GPU in a server.

Most confusion in circulating numbers — including a widely-quoted "3.6M unit backlog" that turns out to be a mislabeled slide about top-4 CSP orders — comes from mixing these units. Every figure below states its unit. Treat any "millions of GPUs" headline without one as unverified.

What has actually shipped

The reconstruction, counting GPU packages, from earnings anchors and supply-chain analysts:

2024: the false start. Blackwell was announced in March 2024; a CoWoS-L packaging rework (warpage, bridge-die redesign — SemiAnalysis, Aug 2024) pushed volume out by months. Morgan Stanley cut its Q4'24 production estimate mid-stream; meaningful revenue only appeared in NVIDIA's fiscal Q4 (Nov'24–Jan'25): $11.0B of Blackwell revenue — "the fastest product ramp in our company's history" (CFO Colette Kress).

2025: the ramp. GB200 NVL72 racks became the volume vehicle (~80% of GB-deployments). Morgan Stanley counted ~27–29k NVL72 racks shipped in CY2025 (~2.0M packages), plus ~1.2M packages as HGX boards and DGX systems. Triangulating Huang's own "3.0M packages through October," Morgan Stanley's rack counts, and Epoch AI's chip-sales model (median 3.4M):

Cumulative Blackwell through Dec 31, 2025: ~3.0–3.2M GPU packages — roughly 2.0M locked inside NVL72 racks, ~1.2M as 8-GPU HGX servers.

2026: acceleration, not transition. A mid-2025 sell-side narrative (JP Morgan: Blackwell falls to 1.8M in 2026 as Rubin ramps) has inverted. TrendForce revised Blackwell's share of NVIDIA's 2026 high-end shipments up from 61% to 71% (April 2026) as Rubin slipped on HBM4 validation; Morgan Stanley now projects 70,000–80,000 NVL72 racks in CY2026 — ~2.6× last year, running at ~8,000 racks/month by June. Data-center revenue confirms the slope: $62.3B (Nov–Jan) → $75.2B (Feb–Apr, +92% YoY) → $89.0B (May–Jul, +117% YoY, reported August 26, "driven by the ramp of our Blackwell Ultra infrastructure"), with the October quarter guided to $108B. In NVIDIA's words, Blackwell 300 was "the fastest product ramp in our company's history." (Yes, they said that twice, about two different Blackwells.)

The monthly tape (Morgan Stanley channel checks): ~6,500 racks in February → ~8,000 in June → an ~8k plateau in July — roughly 40–43k racks in H1 alone — with Hon Hai guiding "high double-digit" sequential growth for Q3. Rough conversion: 2026 is tracking toward 5–6M+ packages for the year.

So where does the total stand today, not last December? Epoch AI's chip-sales model — the only independent public count — puts cumulative Blackwell at 4.8M packages through March 2026 and 5.2M through April. Stack the rack run-rate on top and you get roughly 7–7.5M packages cumulative as of mid-August (our derivation; ~6M of them inside NVL72 racks). The shipped base has more than doubled since New Year — and every number in the next two sections should be read against that August base, not the end-2025 one.

The B300 slice

Blackwell Ultra (B300/GB300, 288GB HBM3e per GPU) went from first commercial deployment (CoreWeave, July 3, 2025) to surpassing GB200 in Blackwell revenue within one quarter: on the November earnings call, Kress reported GB300 "crossed over GB200 and contributed roughly two-thirds of total Blackwell revenue." By the same November quarter it was already, in NVIDIA's written words, the "leading architecture across all customer categories."

Cumulatively, the fleet flipped this year. It was ~75% B200 / 25% B300 at end-2025 — but 2026 production runs 70–80% GB300 (analysts expect ~55k of this year's 70–80k racks to be GB300; Epoch's model counts ~993k B300 packages against ~347k B200 in Q1 alone). By this August, B300-class silicon is the majority of the shipped base: roughly 3.5–4.5M packages, or 50–60% (our derivation). The B300 fleet is no longer the newest sliver of the market — it is the market. It's just not the rentable part.

Where it all goes

Follow the money. The four largest buyers of AI infrastructure spent $413B on capex in 2025 (Amazon $131.8B, Microsoft $118B, Alphabet $91.4B, Meta $72.2B — earnings filings). Their 2026 guidance, as of August: **$720–760B combined** — Amazon raised to ~$220B in July ("higher cost of memory"), Alphabet lifted its ceiling twice to $195–205B, Meta twice to $130–145B. NVIDIA's top five customers alone account for "a little over 50%" of data-center revenue.

Then the anchor deals: OpenAI–NVIDIA at ≥10 GW — NVIDIA's CFO now counts OpenAI's "existing and planned commitments" at ~12 GW, and in August NVIDIA guaranteed up to $105B of land-power-and-shell financing for a 4.25 GW OpenAI campus it says will hold "approximately 1.5 million NVIDIA GPUs" per hardware generation.

The list goes on: OpenAI–AWS at $38B ("hundreds of thousands of GB200/GB300"), with AWS adding "an additional 2 million GPUs" through early 2029; Anthropic at $30B of Azure compute on Grace Blackwell and Vera Rubin; DOE's Solstice at 100,000 Blackwell GPUs. (Unit note: deal figures in these two paragraphs are vendor-reported — gigawatts as power commitments, "GPUs" as the vendors count them; see the unit problem above.) And the sentence that should end any debate about near-term gluts — Amazon's CEO, July 2026: "The lion's share of capacity in 2027 is largely reserved, and we have quite a bit of capacity that's already been reserved for 2028."

Hyperscalers, frontier labs, and sovereigns consume most of their deliveries internally. Some Blackwell does surface publicly — you can rent a GB300 instance from AWS or OCI today — but at $15–18/GPU·hr list it's available, yet not price-competitive with merchant bare metal. NVIDIA's new reporting split (introduced this quarter) makes the shape visible: of Q2's $89B in data-center revenue, $48.7B came from hyperscalers and $40.3B from "AI clouds, industrial & enterprise" — and while the neocloud cohort is growing faster (+138% YoY, on track to exit 2026 with ~8 GW installed, up from ~3 GW), that capacity too is overwhelmingly pre-sold on multi-year contracts before it energizes. What reaches the open market at competitive prices is a far thinner slice.

The rentable sliver

So take the ~7M+ packages shipped through mid-August 2026. What can you — a team that needs 8 to 64 GPUs for a quarter — actually rent?

Start with structure: ~6M of the ~7M sit inside NVL72 racks (up from 2.0M of 3.2M at end-2025 — the rack share is growing), overwhelmingly inside hyperscaler and frontier-lab fleets. The merchant-rentable world lives mostly in the ~1–1.8M HGX-server slice, minus everything pre-sold on multi-year contracts.

Now measure it. On getdeploying's B300 index (Aug 18): 23 providers list B300 at all — fewer than B200 (27), H200 (37), or H100 (53). Of those 23, 15 publish an on-demand price and only 10 show live stock. B300 is the thinnest supply segment in the market. On one of the largest B2B cluster marketplaces, we counted 34 public B300 listings this week: every EU listing carries 5-year terms only with Feb 2027+ delivery; US listings cluster at 36–60 months; listings with any term under 12 months: two. Earliest available date on the entire board: October 8 — one 8-node lot.

Then try to count actual available GPUs rather than listings, and the number collapses further. Vast.ai's marketplace API this week: 7 B300 machines listed — 48 GPUs — zero rentable at the moment of the snapshot. RunPod: out of capacity on both tiers. Shadeform: two clouds carry B300 at all. SF Compute: no B300 market until "this fall." The hyperscalers will sell you GB300 — Azure's ND GB300 v6 and AWS's P6-B300 both went GA in November — but AWS routes on-demand through account managers and Capacity Blocks, Google's A4X Max is reservation-only, and the fleets themselves are spoken for: Microsoft alone contracted ~200,000 GB300s from Nscale, CoreWeave told investors it is "largely sold out of 2026 capacity," and Nebius's August call described its latest capacity auction clearing "15% above the highest price ever charged for Blackwell chips."

That's the actual market for short-term Blackwell. Not 7 million GPUs. Two sub-12-month listings — and 48 marketplace GPUs listed, none immediately rentable at snapshot time.

Why prices rise while shipments boom

Because the sliver, not the shipped base, sets the price — and demand is growing into the sliver faster than it expands.

First, the price signal itself, decomposed — with one honesty note up front: we lack a per-provider November snapshot, so the full 57 points can't be attributed one by one. But from March onward, where archive panels exist, the documented driver is same-provider repricing: Enverge went $5.00→$7.50, Verda $5.74→$7.50, Nebius's own price list moved B300 from $6.10 to $7.85 between May and July — while cheap new entrants roughly offset the hyperscaler listings that joined, leaving index composition a minor term. B200 tells the same story: of twelve providers listed since March, eight raised prices and none cut; Lambda's cluster pricing nearly doubled. One honest caveat in the other direction: the entire repricing wave happened between March and early July — like-for-like prices have been flat for the past six weeks. The market ratcheted up, then held.

The rest of the evidence is behavioral: SemiAnalysis reported in April that on-demand capacity was sold out across all GPU types and new Blackwell deployments were quoting into mid-2026 and beyond; buyers are signing 2027 delivery at 2026 prices; marketplace boards price future delivery with no discount for the wait. H100 one-year contracts — three-year-old silicon — bottomed at $1.70 in October 2025 and have since rebounded 40% as agent workloads ate every idle card; the full curve gets its own section below. When last-generation prices are "melting up" (Latent.Space's phrase), the current generation doesn't crash.

The supply side can't simply flood the market either. TSMC's CoWoS advanced packaging — the bottleneck through which every Blackwell passes — runs at ~70–80k wafers/month in 2025, targeting 130–140k by end-2026, with the entire lineup fully booked and NVIDIA holding roughly half. TSMC's CEO, January 2026: "The capacity is very tight… probably this year, next year, we have to work extremely hard to narrow the gap." Every HBM maker has declared 2026 output sold out, including HBM4 that hasn't fully ramped — and NVIDIA itself just more than doubled its purchase commitments in a single quarter, $119B → $279B, "primarily related to the procurement of memory," while its CFO describes "extreme pricing conditions in memory." The machine is running flat-out, and flat-out is spoken for.

Won't Rubin fix it? Not in this window

Rubin is real — more real than a week ago: full production declared at CES, first VR200 samples shipped in February, and on the August 26 call NVIDIA confirmed production shipments commenced in early August, guiding Rubin to ~20% of data-center revenue already in the October quarter. That guidance makes the analyst consensus for calendar-2026 volume (~250–350k packages; Kuo's bullish rack count implied up to ~500k) look conservative — yet even a doubled figure stays single-digit percent of the shipped base. TrendForce had cut Rubin's 2026 shipment share from 29% to 22% as HBM4 requalification (NVIDIA raised the spec mid-stream), CX8→CX9 networking, and cooling all slipped right; the August ramp says the slippage stopped, not that the volumes are suddenly large.

And Rubin doesn't slot into existing rooms: a VR200 NVL72 rack draws 190–230 kW — roughly double a GB300 rack — is 100% liquid-cooled with no air-cooled variant even for the HGX form factor, and greenfield data centers built for that density take 18–36 months. First-wave allocation goes to hyperscalers; independent clouds are being quoted mid-2027. Huang himself, May 2026: "My sense is that we'll be supply constrained throughout the entire life of Vera Rubin."

Price pressure on Blackwell rentals from Rubin is a 2027–2028 story. In the window that matters for anyone signing a 6–12 month contract today, Rubin adds almost nothing to rentable supply.

The H100 lesson — read it carefully

The standard cautionary tale: H100 rentals peaked above $8/hr in late 2023, spot auctions hit $1–2 within a year, one-year contracts bottomed at $1.70 two years after peak. Cards that sold for $40k trade at $6–15k. If you locked a multi-year rental at peak pricing, you overpaid catastrophically. All true — and it's why we don't sell 5-year terms.

But the second half of the curve is the part nobody quotes: from the October 2025 bottom, H100 contract prices rebounded 40% by March 2026, on-demand sold out again, and the erosion stopped well above zero — because inference and agent workloads found the floor. Silicon depreciates; it doesn't evaporate. The honest model for Blackwell: hold through the current shortage window (prices, lead times, sold-out boards all point at 6–12+ more months), erode as Rubin volume actually lands and 2025's multi-year contracts start expiring (2027 onward), then find a workload floor rather than a cliff.

If anyone offers you 36–60 months of B300 at $12/GPU·hr — now you know exactly whose depreciation risk you'd be holding, and what the last generation did to people who held it.

What we do about it (and what you might)

We built our own economics around this curve: whole 8×B300 nodes on 1–12 month terms, because renters shouldn't carry the multi-year depreciation risk of hardware that history says will reprice. Our reserved 12-month rate ($5.30/GPU·hr) sits at the bottom of the market's published reserved spread ($5.30–7.94) — possible because we buy the supply curve honestly instead of hedging it onto customers. First cluster lands late Q4; the grid is public.

If you're buying compute anywhere: match contract length to your conviction, demand acceptance benchmarks before billing starts, and ask every "great deal" one question — who ends up holding the silicon risk?

Every figure above carries a source — the full fact table with links and methodology notes is published alongside this post: The Blackwell Fact Table. Corrections welcome — that's the point of publishing.

Updated Aug 27 with Q2 FY27 results: revenue $96.2B, data center $89.0B (+117% YoY), October quarter guided to $108B; Vera Rubin production shipments commenced in early August and are guided to ~20% of Q3 data-center revenue. Next checkpoint: the Q3 report in November.