How much VRAM to run Qwen3 Coder 30B-A3B?

About 22 GB atQ4_K_M with an 8K context — fits a RX 7900 XTX. Full breakdown below, or check your exact hardware.

Qwen3 Coder 30B-A3B VRAM by quantisation

QuantisationWeightsTotal (8K ctx)Fits on
Q4_K_M18.6 GB22.4 GBRX 7900 XTX, RTX 4090
Q5_K_M21.7 GB25.8 GBRTX 5090, Radeon AI PRO R9700
Q6_K25.0 GB29.4 GBRTX 5090, Radeon AI PRO R9700
Q8_032.5 GB37.6 GBRTX 6000 Ada, Apple Silicon 64GB unified
FP16 / BF1661.0 GB69.0 GBA100 80GB, RTX PRO 6000 Blackwell

Check your hardware

About Qwen3 Coder 30B-A3B

Qwen3 Coder 30B-A3B is Alibaba's 30.5B-parameter model released in July 2025, with a 256K-token context window. It is a mixture-of-experts model: all 30.5B parameters must sit in memory, but only ~3.3B are active per token, which is what makes it fast for its size. It uses a classic dense-attention design whose KV cache grows linearly with context: its KV cache is about 0.8 GB at an 8K context, 12.9 GB at 128K, and 25.8 GB at the full 256K window (FP16 cache).

For most people Q4_K_M is the sweet spot — the most popular quality/size trade-off — while Q8 is near-lossless if you have the memory. Totals above include the KV cache and a realistic framework overhead, so they are what you should expect to see in practice rather than just the download size. Weight sizes are calibrated against real GGUF files — see themethodology.

Frequently asked questions

How much VRAM does Qwen3 Coder 30B-A3B need?

At Q4_K_M with an 8K context, Qwen3 Coder 30B-A3B needs about 22 GB (weights 19 GB + KV cache + overhead). The smallest common hardware that fits is a RX 7900 XTX.

Can an RTX 4090 (24GB) run Qwen3 Coder 30B-A3B?

Yes. An RTX 4090's 24 GB runs Qwen3 Coder 30B-A3B at Q4_K_M (about 22 GB at 8K context).

Can a Mac run Qwen3 Coder 30B-A3B?

Yes — Apple Silicon with 32 GB of unified memory or more (macOS lets the GPU use ~75% of it, ~24 GB) runs Qwen3 Coder 30B-A3B at Q4_K_M.

Related

VRAM calculator for any model ·Token counter

Last updated 2026-08-03. Architecture figures from the model's published config.json; see themethodology.