Three open-weight models. For each one, we show you the math and the minimum hardware to run it.
Parameters (billions) × 1 GB = base memoryBase × 1.20 = total minimum (the 20% covers activations, KV cache, and working space)A capable, compact general-purpose chat and reasoning model. Fast to serve, cheap to run. Ideal for embedded applications, internal tools, and cost-sensitive production.
The full-power GLM 5.2. Competitive on benchmarks with leading proprietary models. Strong multilingual reasoning, advanced coding, enterprise automation.
A massive Mixture-of-Experts coding model. State-of-the-art on code generation and agentic tasks. This is where serious infrastructure investment pays off.