On an RTX 5070 Ti (16 GB), Ministral 3 14B in Q4_K_M needs about 9,2 Go of video memory: it should fit. Estimated generation speed: ≈ 88,9 tokens/s.
Not measured yet: the figures below are estimates, computed from the models already measured on this card. Measure my card
≈ 88,9 tok/s
Generation (estimated)
≥ 9,2 Go
Recommended VRAM (16 GB avail.)
≈ 179 W
Power (ref. model)
0,14 €/M tok
Electricity / M tokens · 0,25 €/kWh
Calculation: DeepSeek R1 Distill Qwen 14B measured at 81,5 tok/s on this card, 8,4 Go read per token → × 8,4 Go ÷ 7,7 Go ≈ 88,9 tok/s (generation is memory-bound: speed follows the amount of weights read).
What about a longer context?
Context
KV cache
Estimated VRAM
Fits in 16 GB?
Estimated speed
4 096
0,6 Go
8,8 Go
yes
≈ 83,2 tok/s
8 192
1,3 Go
9,4 Go
yes
≈ 77,3 tok/s
16 384
2,5 Go
10,7 Go
yes
≈ 67,7 tok/s
32 768
5,0 Go
13,2 Go
yes
≈ 54,2 tok/s
65 536
10,0 Go
18,2 Go
no
—
131 072
20,0 Go
28,2 Go
no
—
Estimated VRAM = file + KV cache (full-attention layers, f16) + ~0,5 Go of buffers, excluding memory used by the system — hence the slightly higher “recommended VRAM”, which keeps that margin. Speed ≈ estimated speed × weights read ÷ (weights read + added KV cache).
Estimates based on real measurements of other models on this card and on the model's public data (Mistral). As soon as a measurement exists, this page shows the real figures.