Leaderboard

NVIDIA

RTX 3060

260 € checked on 04/08/26

12 Go

VRAM

48,0 tok/s

Best generation

0,31 tok/s/W

Efficiency (ref.)

0,22 €/M tok

Elec. cost / M tokens (ref.) · 0,25 €/kWh

What this card can run

Llama 3.2 3B Instruct Q4_K_M

2.02 Go

Not measured

Qwen2.5 7B Instruct Q4_K_M

4.68 Go

Not measured

Meta Llama 3.1 8B Instruct Q4_K_M★

4.92 Go · 48,0 tok/s measured

Comfortable

Qwen2.5 14B Instruct Q4_K_M

8.99 Go

Not measured

DeepSeek R1 Distill Qwen 14B Q4_K_M

8.99 Go

Not measured

Phi-4 Q4_K_M

9.05 Go

Not measured

Smooth ≥ 60 tok/s (faster than you read) · Comfortable 30-60 · Slow < 30 · "Doesn't fit" = model size + 2 GB headroom > VRAM (CPU offloading is excluded from the protocol: the numbers would be meaningless).

Measurements by model

48,0 tok/s

Generation (avg)

0,31 tok/s/W

Efficiency

0,22 €/M tok

Electricity cost

1

Runs

Cost = electricity only (€0.25/kWh), to generate 1 million tokens at the measured speed — the billing unit of cloud APIs, and the comparison is deliberate.

Evolution over time

Each dot = one run (llama.cpp version on hover). History builds up by re-benching the same card across versions — drivers and llama.cpp keep moving the numbers.

DateVersionGenerationPromptVRAM peakMax tempPowerContextSource
04/08—48,01 5005,9 Go68 °C155 W—demo

VRAM used = peak system GPU memory during the run (includes the OS, roughly 0.5–1 GB more than the model alone).