The Cheapest Honest Build

You want one big open model running in production. We'll show you every viable option — and be straight with you about the tradeoffs.

How this works
Each build below runs a real AI model on real hardware. We show you the model's memory requirement, the hardware that meets it, the power it draws, and the total price. Pick the build that fits — then request a quote.
Option 1

The Lean Start

Run GLM 5.1 — one card, workstation-grade
$5,000
total build price

GLM 5.1 on a single RTX 4090 mounted in a small-form-factor workstation. No datacenter, no rack, no cooling contract. Fits under a desk. Serious enough for a real product.

The model you're running
GLM 5.1
Z.ai · MIT · via SiliconFlow
9B parameters
9B × 1 GB9 GB
+20% working room+2 GB
Minimum VRAM11 GB
What you're buying
NVIDIA GeForce RTX 4090
24 GB GDDR6X · 450 W each · PCIe 4.0 x16
$2,399
Power reality
600 W
Total draw
=
0.5
Average homes
·
0.2 hrs
To fill an EV battery
Half a typical home. Runs on a standard circuit.
Option 2

The Real Start

Run GLM 5.2 MAX — two H100s, server-grade
$72,000
total build price

GLM 5.2 MAX spread across 2 H100 PCIe cards in a dual-GPU server. 160 GB combined. Serious model, serious throughput. Runs in a standard server rack.

The model you're running
GLM 5.2 (MAX)
Z.ai · MIT · via SiliconFlow
130B parameters
130B × 1 GB130 GB
+20% working room+26 GB
Minimum VRAM156 GB
What you're buying
NVIDIA H100 PCIe 80 GB
80 GB HBM3 · 350 W each · PCIe 5.0 x16
$60,000
Combined VRAM160 GB Combined Power1,800 W
Power reality
1,800 W
Total draw
=
1.5
Average homes
·
0.5 hrs
To fill an EV battery
1.5 homes. One 30-amp rack PDU is enough.
Need more capacity?

The startup builds above cover one model, one deployment. If you need to serve hundreds of users or run a larger model, the mid-size path starts where this one ends.

Mid-size path →