Alibaba

Qwen3

reasoningMoE

Pick a size: the video memory shown is the Q4_K_M file (the protocol's quantisation) plus about 1.5 GB for a 4,000-token context.

0.6B

8 GB card
Parameters
0,8B
Q4_K_M file
0,5 Go
Recommended VRAM
≥ 2,0 Go
Max context
40 960

Measured on 1 card · 728 tok/s (RTX 5070 Ti)

1.7B

8 GB card
Parameters
2B
Q4_K_M file
1,2 Go
Recommended VRAM
≥ 2,7 Go
Max context
40 960

Measured on 1 card · 456 tok/s (RTX 5070 Ti)

4B

8 GB card
Parameters
4B
Q4_K_M file
2,3 Go
Recommended VRAM
≥ 3,8 Go
Max context
40 960

Measured on 1 card · 225 tok/s (RTX 5070 Ti)

8B

8 GB card
Parameters
8,2B
Q4_K_M file
4,7 Go
Recommended VRAM
≥ 6,2 Go
Max context
40 960

Measured on 1 card · 146 tok/s (RTX 5070 Ti)

14B

12 GB card
Parameters
14,8B
Q4_K_M file
8,4 Go
Recommended VRAM
≥ 9,9 Go
Max context
40 960

Measured on 1 card · 84,8 tok/s (RTX 5070 Ti)

32B

24 GB card
Parameters
32,8B
Q4_K_M file
18,4 Go
Recommended VRAM
≥ 19,9 Go
Max context
40 960

Not measured yet

30B-A3B

24 GB card
Parameters
30,5B · 3B active
Q4_K_M file
17,4 Go
Recommended VRAM
≥ 18,9 Go
Max context
40 960

Not measured yet

235B-A22B

> 32 GB
Parameters
235,1B · 22B active
Q4_K_M file
132,4 Go
Recommended VRAM
≥ 133,9 Go
Max context
40 960

Not measured yet

MoE (“30B-A3B”): all the memory of a 30B, but the speed of a model with 3B active parameters. “> 32 GB”: several cards, or part of the model in system memory (much slower).