# GLM-4.5-Air: RAM and VRAM requirements

> GLM-4.5-Air is a 106B GLM model (Mixture-of-Experts, 12B active per token). At Q4_K_M it needs about **71.8 GB** to run and fits **4 of 40** tracked devices. Minimum to run: Apple M4 Max (128GB).

Last validated: 2026-08-03. Sources: Ollama, HuggingFace GGUF repos, vendor specs.

## Memory by quantization
| Quant | On disk | To run (4k context) |
| --- | --- | --- |
| Q4_K_M | 68.45 GB | ~71.8 GB |
| Q8_0 | 109.39 GB | ~112.7 GB |

Memory = weights + KV cache + ~0.8 GB runtime overhead, and varies ±15% with context length.

## Will it run on my device?
- **Apple M1 (8GB)** (8 GB): No, not enough memory
- **iPhone 17** (8 GB): No, not enough memory
- **iPhone 17 Pro** (12 GB): No, not enough memory
- **16GB RAM Laptop (CPU/iGPU only)** (16 GB): No, not enough memory
- **Google Pixel 10 Pro** (16 GB): No, not enough memory
- **Nvidia GeForce RTX 3090 (24GB)** (24 GB): No, not enough memory
- **Apple M4 Pro (48GB)** (48 GB): No, not enough memory
- **Apple M3 Ultra (256GB)** (256 GB): Yes, it runs (room for Q8_0)

Full table of all 40 devices: https://localmodel.run/model/glm-4.5-air

## How to run
Use LM Studio (Mac/Windows) or Ollama / vLLM (Linux).

## Details
- Parameters: 106B (MoE, 12B active per token)
- Default context: 128k tokens
- License: MIT (commercial use: yes)
- Released: 2025-07
- HuggingFace: 441,341 downloads/mo, 625 likes

## FAQ
### How much VRAM or RAM does GLM-4.5-Air need?
About 71.8 GB at Q4_K_M (weights 68.45 GB + KV cache + overhead) at a 4k context. Budget ~112.7 GB for Q8_0.
### Can GLM-4.5-Air run on a laptop?
GLM-4.5-Air is large; you need a high-memory Mac or a 24 GB+ GPU at Q4_K_M.
### Can I use GLM-4.5-Air commercially?
Yes, MIT permits commercial use.

Sources: https://huggingface.co/zai-org/GLM-4.5-Air, https://huggingface.co/bartowski/zai-org_GLM-4.5-Air-GGUF, https://artificialanalysis.ai/models/glm-4-5-air, https://llm-stats.com/models/glm-4.5-air, https://docs.z.ai/guides/llm/glm-4.5
More: https://localmodel.run/model/glm-4.5-air