DeepSeek

R1 Distill

reasoning

Pick a size: the video memory shown is the Q4_K_M file (the protocol's quantisation) plus about 1.5 GB for a 4,000-token context.

1.5B Qwen

8 GB card
Parameters
1,8B
Q4_K_M file
1,0 Go
Recommended VRAM
≥ 2,5 Go
Max context
131 072

Not measured yet

7B Qwen

8 GB card
Parameters
7,6B
Q4_K_M file
4,4 Go
Recommended VRAM
≥ 5,9 Go
Max context
131 072

Not measured yet

8B Llama

8 GB card
Parameters
8B
Q4_K_M file
4,6 Go
Recommended VRAM
≥ 6,1 Go
Max context
131 072

Not measured yet

14B Qwen

12 GB card
Parameters
14,8B
Q4_K_M file
8,4 Go
Recommended VRAM
≥ 9,9 Go
Max context
131 072

Measured on 1 card · 81,5 tok/s (RTX 5070 Ti)

32B Qwen

24 GB card
Parameters
32,8B
Q4_K_M file
18,5 Go
Recommended VRAM
≥ 20,0 Go
Max context
131 072

Not measured yet

70B Llama

> 32 GB
Parameters
70,6B
Q4_K_M file
39,6 Go
Recommended VRAM
≥ 41,1 Go
Max context
131 072

Not measured yet

MoE (“30B-A3B”): all the memory of a 30B, but the speed of a model with 3B active parameters. “> 32 GB”: several cards, or part of the model in system memory (much slower).