Everyone is rushing to run the big open AI models, and they all need the same thing: serious NVIDIA hardware. It's expensive and the spec sheets are unreadable. This store sells the real thing — from a $1,999 desktop card to a $6.5 million rack — and explains every number in words a human can feel.
The GPU's own super-fast storage. An AI model must fit entirely inside GPU memory to run — like a book that must lie open on the desk, not on a shelf. Rule of thumb: 1 billion parameters ≈ 1 GB, plus 20% working room. This is the number that decides which models you can run at all.
How much electricity the machine eats every second it's on. It decides your monthly bill and whether your building can even feed the machine — a normal wall outlet gives ~1,800 W; a big rack needs 135,000 W. For scale: an average home draws about 1,200 W around the clock, and a typical electric-car battery holds 90 kWh.
What you actually pay up front. In AI hardware you are mostly paying for memory, and for how fast that memory is — datacenter cards cost more per GB because their HBM memory is ~5× faster than desktop memory, which is what keeps answers fast when real users are waiting.
Also on spec sheets: CPU and RAM — the machine's general-purpose brain and ordinary memory. They load data and run your website, but the AI model itself lives and works in GPU memory. For running models, GPU memory is the number to watch.
Real products, real prices, real watts. Every card and system below is a current NVIDIA product; rack systems are sold through OEM partners, so those prices are the reported market figures.
The biggest gaming/creator card. The cheapest honest way to experiment with small open models (up to ~26B parameters) at home.
When it matters: Matters when you're a developer prototyping on your desk. 32 GB fits small models only — a 284B model will never fit on one, or even on eight of these stacked together in a normal PC.
The professional big-memory card. 96 GB is 3× the RTX 5090 for ~4× the price — you're paying for memory, and for AI that's the right trade.
When it matters: Matters when one card must hold a serious chunk of a model. Four of these in one server is the cheapest honest home for a ~284B model.
A versatile datacenter card for inference, graphics and video. Lower power draw per card makes it popular in ordinary server rooms.
When it matters: Matters when your rack has limited power/cooling: 350 W per card is almost half an RTX PRO 6000. But 48 GB means many cards for big models.
A Hopper datacenter GPU with 141 GB of HBM3e — memory that is roughly 5× faster than the GDDR7 in desktop cards. Speed of memory decides how fast a model can actually talk.
When it matters: Matters when you serve real user traffic: HBM bandwidth is what keeps response times low when a big model generates text token by token.
NVIDIA's turnkey AI server: 8 Blackwell GPUs joined by NVLink so they behave like one giant 1.4 TB GPU. This is the standard building block of serious AI companies.
When it matters: Matters when a model is too big for any single card: NVLink lets the 8 GPUs share the model at chip speed instead of network speed.
A full liquid-cooled rack: 72 Blackwell GPUs wired into a single NVLink domain — the whole rack acts like one enormous GPU with 13.4 TB of memory.
When it matters: Matters when you run trillion-parameter models or serve thousands of concurrent users. Requires real datacenter power and liquid cooling.
The current top of the line: 72 Blackwell Ultra GPUs, ~21 TB of HBM3e in one rack. The machine the AI gold rush is actually fought with.
When it matters: Matters when you are a frontier AI lab or a cloud selling AI capacity. One rack draws as much power as about 112 homes — plan the building first.
Three big open-source models from the world agent leaderboard (filtered to Open Source, MIT license). The memory rule is simple and we show the math: parameters in billions × 1 GB, plus 20% working room. The 20% covers the "scratch paper" the model needs while it thinks — your conversation, its intermediate calculations.
Only 13B of the 284B parameters 'wake up' per word (mixture-of-experts), so it runs fast — but ALL 284B must sit in GPU memory at once.
At 744B parameters this no longer fits any workstation — you are shopping for a datacenter-class 8-GPU server.
1.6 trillion parameters. No single machine sold today holds 1,920 GB — this is where clustering stops being optional.
You want to run one big open model behind your product, as cheaply as it can honestly be run. Cheapest honest choice of our three models: DeepSeek V4 Flash (284B → needs 340.8 GB).
Why not stack RTX 5090s? Eleven desktop cards can't live in one PC, and split across several PCs they'd talk over a slow network — see the cluster truth below. Four big-memory cards in one box is the honest floor.
More users, more traffic, and you need headroom to scale. You want a machine that runs GLM 5.2 (Max) (744B → needs 892.8 GB) today and still has room for tomorrow.
The headroom is the point: room for longer conversations, more simultaneous users, or a second model. When you outgrow it, add a second DGX B200 and cluster them — that's the Frontier Cluster below, and it unlocks DeepSeek V4 Pro (1.6T).
The most expensive mistake in AI hardware is buying a machine you only needed for a month. Pick a model, slide how many GPU-hours you actually need, and watch the honest answer update live. Assumptions, stated plainly: cloud prices are typical on-demand market rates; electricity costs $0.15 per kWh.
Electricity, cooling and hosting are inside the hourly price — you pay the meter and walk away when you're done.
A cluster is several machines wired together so closely that they can work on one job as if they were a single computer.
DeepSeek V4 Pro needs 1,920 GB. No single machine sold today holds that, so we combine two DGX B200s:
Inside one server, GPUs talk over NVLink — a private highway moving ~1,800 GB/s. Between servers, InfiniBand moves ~50–100 GB/s. Both are fast enough that a model split across the machines still answers quickly. That's what you're paying for in DGX systems and NVL72 racks: the wiring, not just the chips.
Regular Ethernet moves ~1–10 GB/s — hundreds of times slower than NVLink. Split a big model across gaming PCs on an office network and the GPUs spend their time waiting for the network, not computing. Ten RTX 5090s look like 320 GB on paper, but as a "cluster" they crawl. Honest advice: desktop cards are for one-box work only.
Yardsticks: an average home draws about 1,200 W around the clock; a typical electric-car battery holds about 90 kWh (so a machine drawing 90,000 W drains one full EV battery every hour).
| Build | Memory | Power | ≈ homes | EV battery lasts | Total price |
|---|---|---|---|---|---|
| Startup Build — one server, one model 1 server with 4 × RTX PRO 6000 Blackwell |
384 GB | 3,200 W | 2.7 homes | 28.1 h | $48,260 |
| Growing Build — NVIDIA DGX B200 1 × DGX B200 (8 NVLink-connected GPUs) |
1,440 GB | 14,300 W | 11.9 homes | 6.3 h | $515,000 |
| Frontier Cluster — 2 × DGX B200 2 × DGX B200 joined by InfiniBand |
2,880 GB | 28,600 W | 23.8 homes | 3.1 h | $1,030,000 |
| Datacenter Rack — GB300 NVL72 1 rack: 72 Blackwell Ultra GPUs, one NVLink domain |
20,736 GB | 135,000 W | 112.5 homes | 0.7 h | $6,500,000 |
Read the last row out loud: the GB300 NVL72 rack draws 135,000 W — as much power as about 112 homes, and it drains a full electric-car battery every 40 minutes. This is why datacenters are built next to power plants.
Leave your name and contact info and which build you want — we'll get back to you with a formal quote. You'll get a request number right away.