Serious Capacity

More users. More traffic. Real redundancy. Room to grow. These builds are for teams who've outgrown the proof-of-concept stage.

Where the startup path left off
The startup path ended at 2 H100s ($72,000) running GLM 5.2 MAX. The builds below are for when that's not enough — more users, larger models, or both. Every build here is production-grade with real cluster options.
Option 1

Scale Ready

GLM 5.2 MAX — four H100s, production capacity
$138,000
total build price

4× H100 PCIe in a quad-GPU server. Run GLM 5.2 MAX with 2× redundancy — serve double the users, keep one copy in reserve. First real production-grade build.

The model you're running
GLM 5.2 (MAX)
Z.ai · MIT · via SiliconFlow
130B parameters
130B × 1 GB130 GB
+20% working room+26 GB
Minimum VRAM156 GB
What you're buying
NVIDIA H100 PCIe 80 GB
80 GB HBM3 · 350 W each · PCIe 5.0 x16
$120,000
Combined VRAM320 GB Combined Power3,000 W Total Price$138,000
Power reality
3,000 W
Total draw
=
2.5
Average homes
·
0.8 hrs
To fill an EV battery
2.5 homes. Needs a dedicated rack PDU circuit.
Option 2

The DGX Stack

Kimi K2.7 Code — one DGX H200, full model
$465,000
total build price

One DGX H200. 1.1 TB of NVSwitch-unified GPU memory. Runs Kimi K2.7 Code in full. The leap from loose PCIe cards to real fabric — all 8 GPUs act as one.

DGX system — 8 GPUs wired by NVSwitch fabric inside one box. Not a cluster, but acts like one.
The model you're running
Kimi K2.7 Code
Moonshot · Modified MIT · via Fireworks
671B parameters
671B × 1 GB671 GB
+20% working room+135 GB
Minimum VRAM806 GB
What you're buying
NVIDIA DGX H200
1,128 GB (8 × 141 GB HBM3e) · 10,200 W each · 8× SXM5, NVSwitch
$450,000
Combined VRAM1,128 GB Combined Power11,500 W Total Price$465,000
Power reality
11,500 W
Total draw
=
9.6
Average homes
·
3.1 hrs
To fill an EV battery
9.6 homes. Requires datacenter-grade PDU and cooling.
Option 3

The Cluster

Kimi K2.7 Code at scale — four DGX H200 nodes
$1,900,000
total build price

4 DGX H200 nodes wired by InfiniBand. 4.5 TB GPU memory, true cluster scheduling, real hardware failover. Serve hundreds of concurrent users on the largest open models.

This is a real cluster — multiple machines networked together. Learn what that means →
The model you're running
Kimi K2.7 Code
Moonshot · Modified MIT · via Fireworks
671B parameters
671B × 1 GB671 GB
+20% working room+135 GB
Minimum VRAM806 GB
What you're buying
NVIDIA DGX H200
1,128 GB (8 × 141 GB HBM3e) · 10,200 W each · 8× SXM5, NVSwitch
$1,800,000
Combined VRAM4,512 GB Combined Power46,000 W Total Price$1,900,000
Power reality
46,000 W
Total draw
=
38.3
Average homes
·
12.3 hrs
To fill an EV battery
38 homes. Full datacenter row. Needs facility planning.
Need a full rack?

The builds above top out at 4 DGX H200 nodes. For a full SuperPOD or the $6.5M datacenter rack, reach out directly for a custom quote.

See full rack products → Request custom quote