On an RTX 4070 Ti (12 GB), Phi-3.5 mini 3.8B in Q4_K_M needs about 3,7 Go of video memory: it should fit. Estimated generation speed: ≈ 190 tokens/s.
Not measured yet: the figures below are estimates, computed from the models already measured on this card. Measure my card
≈ 190 tok/s
Generation (estimated)
≥ 3,7 Go
Recommended VRAM (12 GB avail.)
≈ 106 W
Power (ref. model)
0,04 €/M tok
Electricity / M tokens · 0,25 €/kWh
Calculation: Gemma 3n E2B measured at 163 tok/s on this card, 2,6 Go read per token → × 2,6 Go ÷ 2,2 Go ≈ 190 tok/s (generation is memory-bound: speed follows the amount of weights read).
What about a longer context?
Context
KV cache
Estimated VRAM
Fits in 12 GB?
Estimated speed
4 096
1,5 Go
4,2 Go
yes
≈ 121 tok/s
8 192
3,0 Go
5,7 Go
yes
≈ 84,6 tok/s
16 384
6,0 Go
8,7 Go
yes
≈ 52,8 tok/s
32 768
12,0 Go
14,7 Go
no
—
65 536
24,0 Go
26,7 Go
no
—
131 072
48,0 Go
50,7 Go
no
—
Estimated VRAM = file + KV cache (full-attention layers, f16) + ~0,5 Go of buffers, excluding memory used by the system — hence the slightly higher “recommended VRAM”, which keeps that margin. Speed ≈ estimated speed × weights read ÷ (weights read + added KV cache).
Estimates based on real measurements of other models on this card and on the model's public data (Microsoft Phi). As soon as a measurement exists, this page shows the real figures.