Leaderboard

NVIDIA

RTX 4090

24 Go

VRAM

132 tok/s

Best generation

0,46 tok/s/W

Efficiency (ref.)

0,15 €/M tok

Elec. cost / M tokens (ref.) · 0,25 €/kWh

What this card can run

Llama 3.2 3B Instruct Q4_K_M

2.02 Go

Not measured

Qwen2.5 7B Instruct Q4_K_M

4.68 Go

Not measured

Meta Llama 3.1 8B Instruct Q4_K_M★

4.92 Go · 132 tok/s measured

Smooth

Qwen2.5 14B Instruct Q4_K_M

8.99 Go · 78,0 tok/s measured

Smooth

DeepSeek R1 Distill Qwen 14B Q4_K_M

8.99 Go

Not measured

Phi-4 Q4_K_M

9.05 Go

Not measured

Smooth ≥ 60 tok/s (faster than you read) · Comfortable 30-60 · Slow < 30 · "Doesn't fit" = model size + 2 GB headroom > VRAM (CPU offloading is excluded from the protocol: the numbers would be meaningless).

Measurements by model

132 tok/s

Generation (avg)

0,46 tok/s/W

Efficiency

0,15 €/M tok

Electricity cost

1

Runs

Cost = electricity only (€0.25/kWh), to generate 1 million tokens at the measured speed — the billing unit of cloud APIs, and the comparison is deliberate.

Evolution over time

Each dot = one run (llama.cpp version on hover). History builds up by re-benching the same card across versions — drivers and llama.cpp keep moving the numbers.

DateVersionGenerationPromptVRAM peakMax tempPowerContextSource
04/08—1325 9006,2 Go66 °C285 W—demo

VRAM used = peak system GPU memory during the run (includes the OS, roughly 0.5–1 GB more than the model alone).