DeepSeek
R1 Distill
reasoning
Pick a size: the video memory shown is the Q4_K_M file (the protocol's quantisation) plus about 1.5 GB for a 4,000-token context.
1.5B Qwen
8 GB card- Parameters
- 1,8B
- Q4_K_M file
- 1,0 Go
- Recommended VRAM
- ≥ 2,5 Go
- Max context
- 131 072
Not measured yet
7B Qwen
8 GB card- Parameters
- 7,6B
- Q4_K_M file
- 4,4 Go
- Recommended VRAM
- ≥ 5,9 Go
- Max context
- 131 072
Not measured yet
8B Llama
8 GB card- Parameters
- 8B
- Q4_K_M file
- 4,6 Go
- Recommended VRAM
- ≥ 6,1 Go
- Max context
- 131 072
Not measured yet
14B Qwen
12 GB card- Parameters
- 14,8B
- Q4_K_M file
- 8,4 Go
- Recommended VRAM
- ≥ 9,9 Go
- Max context
- 131 072
Measured on 1 card · 81,5 tok/s (RTX 5070 Ti)
32B Qwen
24 GB card- Parameters
- 32,8B
- Q4_K_M file
- 18,5 Go
- Recommended VRAM
- ≥ 20,0 Go
- Max context
- 131 072
Not measured yet
70B Llama
> 32 GB- Parameters
- 70,6B
- Q4_K_M file
- 39,6 Go
- Recommended VRAM
- ≥ 41,1 Go
- Max context
- 131 072
Not measured yet
MoE (“30B-A3B”): all the memory of a 30B, but the speed of a model with 3B active parameters. “> 32 GB”: several cards, or part of the model in system memory (much slower).