AI Models

Three open-weight models. For each one, we show you the math and the minimum hardware to run it.

How the memory math works
Every model needs to live entirely in GPU memory (VRAM) to run at reasonable speed. The formula:
Parameters (billions) × 1 GB = base memory
Base × 1.20 = total minimum (the 20% covers activations, KV cache, and working space)

This assumes INT8 quantization — the standard for production inference. If you run FP16, double the numbers.

GLM 5.1

Z.ai MIT via SiliconFlow
9B
parameters

A capable, compact general-purpose chat and reasoning model. Fast to serve, cheap to run. Ideal for embedded applications, internal tools, and cost-sensitive production.

VRAM Usage on Minimum Hardware (24 GB GDDR6X)
9 GB
+2 GB
13 GB free
Base — 9 GB Working room — +2 GB Headroom — 13 GB free

Memory Math

Parameters
9B params × 1 GB
9 GB
Working room (+20%)
activations + KV cache + buffers
+ 2 GB
Minimum VRAM
to run at all
11 GB
FP16 (full precision) would require ~22 GB. The numbers above are for INT8 — the standard for production serving.

Minimum Hardware

NVIDIA GeForce RTX 4090
24 GB GDDR6X available — 13 GB headroom
$2,399
450 W — 0.4 homes
Request This Build

Use Cases

  • Chat & customer support
  • Document summarization
  • Code assistance
  • Internal tools

Hardware that can run GLM 5.1

NVIDIA GeForce RTX 4090
24 GB GDDR6X
46% used
$2,399
NVIDIA RTX 6000 Ada
48 GB GDDR6 ECC
23% used
$6,800
NVIDIA L40S
48 GB GDDR6 ECC
23% used
$22,000
NVIDIA H100 PCIe 80 GB
80 GB HBM3
14% used
$30,000
NVIDIA H100 SXM5 80 GB
80 GB HBM3e
14% used
$40,000
NVIDIA H200 SXM5 141 GB
141 GB HBM3e
8% used
$45,000
NVIDIA H100 NVL
188 GB HBM3 (2 × 94 GB, NVLink)
6% used
$65,000
NVIDIA DGX H100
640 GB (8 × 80 GB HBM3e)
2% used
$300,000
NVIDIA DGX H200
1,128 GB (8 × 141 GB HBM3e)
1% used
$450,000
NVIDIA DGX SuperPOD H200
11,280 GB (80 × H200 141 GB)
0% used
$3,500,000
NVIDIA MGX Full Datacenter Rack
22,560 GB (160 × H200 141 GB)
0% used
$6,500,000

GLM 5.2 (MAX)

Z.ai MIT via SiliconFlow
130B
parameters

The full-power GLM 5.2. Competitive on benchmarks with leading proprietary models. Strong multilingual reasoning, advanced coding, enterprise automation.

VRAM Usage on Minimum Hardware (188 GB HBM3 (2 × 94 GB, NVLink))
130 GB
+26 GB
32 GB free
Base — 130 GB Working room — +26 GB Headroom — 32 GB free

Memory Math

Parameters
130B params × 1 GB
130 GB
Working room (+20%)
activations + KV cache + buffers
+ 26 GB
Minimum VRAM
to run at all
156 GB
FP16 (full precision) would require ~312 GB. The numbers above are for INT8 — the standard for production serving.

Minimum Hardware

NVIDIA H100 NVL
188 GB HBM3 (2 × 94 GB, NVLink) available — 32 GB headroom
$65,000
600 W — 0.5 homes
Request This Build

Use Cases

  • Complex reasoning
  • Multi-language support
  • Advanced coding
  • Enterprise automation

Hardware that can run GLM 5.2 (MAX)

NVIDIA H100 NVL
188 GB HBM3 (2 × 94 GB, NVLink)
83% used
$65,000
NVIDIA DGX H100
640 GB (8 × 80 GB HBM3e)
24% used
$300,000
NVIDIA DGX H200
1,128 GB (8 × 141 GB HBM3e)
14% used
$450,000
NVIDIA DGX SuperPOD H200
11,280 GB (80 × H200 141 GB)
1% used
$3,500,000
NVIDIA MGX Full Datacenter Rack
22,560 GB (160 × H200 141 GB)
1% used
$6,500,000

Kimi K2.7 Code

Moonshot Modified MIT via Fireworks
671B
parameters

A massive Mixture-of-Experts coding model. State-of-the-art on code generation and agentic tasks. This is where serious infrastructure investment pays off.

VRAM Usage on Minimum Hardware (1,128 GB (8 × 141 GB HBM3e))
671 GB
+135 GB
322 GB free
Base — 671 GB Working room — +135 GB Headroom — 322 GB free

Memory Math

Parameters
671B params × 1 GB
671 GB
Working room (+20%)
activations + KV cache + buffers
+ 135 GB
Minimum VRAM
to run at all
806 GB
FP16 (full precision) would require ~1612 GB. The numbers above are for INT8 — the standard for production serving.

Minimum Hardware

NVIDIA DGX H200
1,128 GB (8 × 141 GB HBM3e) available — 322 GB headroom
$450,000
10,200 W — 8.5 homes
Request This Build

Use Cases

  • Large-scale code generation
  • Autonomous coding agents
  • Complex system design
  • R&D pipelines

Hardware that can run Kimi K2.7 Code

NVIDIA DGX H200
1,128 GB (8 × 141 GB HBM3e)
71% used
$450,000
NVIDIA DGX SuperPOD H200
11,280 GB (80 × H200 141 GB)
7% used
$3,500,000
NVIDIA MGX Full Datacenter Rack
22,560 GB (160 × H200 141 GB)
4% used
$6,500,000