Back to Calculator Deploy on RunPodDeploy Now
GLM-5.1
Z.ai next-gen flagship for agentic engineering. 744B MoE, 40B active. MIT licensed. #1 open-weight model on SWE-Bench Pro as of April 2026. Trained on Huawei Ascend chips.
Specifications
SourceArchitectureTEXT
Parameters744B
Familyglm
VRAM (Q4)372.0G
MoE: 40B active.
codingagentsreasoning
Run in the Cloud
This model requires enterprise-grade VRAM. Rent GPUs on RunPod and start generating.
Instant Cloud GPUs
Running out of VRAM? Rent a high-end H100 or RTX 4090 on RunPod and deploy in seconds.
Quantization Estimates
| Format | VRAM Need | Tier |
|---|---|---|
| FP16 | 1488.0 GB | Full Precision |
| Q8_0 | 744.0 GB | High |
| Q6_K | 632.4 GB | Excellent |
| Q5_K_M | 520.8 GB | Great |
| Q4_K_M | 372.0 GB | Sweet Spot |
| Q2_K | 223.2 GB | Emergency |
Share this Model
Send these specs directly to your community.