NVIDIA
RTX 4070 Ti
927 € checked on 04/08/26
12 Go
VRAM
187 tok/s
Best generation
0,56 tok/s/W
Efficiency (ref.)
0,12 €/M tok
Elec. cost / M tokens (ref.) · 0,25 €/kWh
What this card can run
Llama 3.2 3B Instruct Q4_K_M
2.02 Go · 186 tok/s measured
Qwen2.5 7B Instruct Q4_K_M
4.68 Go · 94,9 tok/s measured
Meta Llama 3.1 8B Instruct Q4_K_M★
4.92 Go · 89,6 tok/s measured
Qwen2.5 14B Instruct Q4_K_M
8.99 Go · 49,2 tok/s measured
DeepSeek R1 Distill Qwen 14B Q4_K_M
8.99 Go · 49,3 tok/s measured
Phi-4 Q4_K_M
9.05 Go · 48,8 tok/s measured
Smooth ≥ 60 tok/s (faster than you read) · Comfortable 30-60 · Slow < 30 · "Doesn't fit" = model size + 2 GB headroom > VRAM (CPU offloading is excluded from the protocol: the numbers would be meaningless).
Measurements by model
186 tok/s
Generation (avg)
1,56 tok/s/W
Efficiency
0,04 €/M tok
Electricity cost
2
Runs
Cost = electricity only (€0.25/kWh), to generate 1 million tokens at the measured speed — the billing unit of cloud APIs, and the comparison is deliberate.
Evolution over time
Each dot = one run (llama.cpp version on hover). History builds up by re-benching the same card across versions — drivers and llama.cpp keep moving the numbers.
| Date | Version | Generation | Prompt | VRAM peak | Max temp | Power | Context | Source |
|---|---|---|---|---|---|---|---|---|
| 08/08 | llama.cpp 96278e39f | 184 | 12 421 | 3,0 Go | 44 °C | 119 W | — | measured |
| 04/08 | llama.cpp 96278e39f | 187 | 13 074 | 3,2 Go | 48 °C | 115 W | — | measured |
VRAM used = peak system GPU memory during the run (includes the OS, roughly 0.5–1 GB more than the model alone).