Configurator

Build your setup, without overflowing

Pick your card, the memory to leave to the system, the model and the context size: we check that everything fits in video memory and estimate the speed.

Speeds at context 0 are measured. Anything depending on context or concurrency is an estimate: every formula is shown, nothing is written to the database.

Video memory of the RTX 4090

11,0 Go used of 23,0 Go usable (24 GB)

12,0 Go left to add a model or grow a context.

▮ model (measured)▨ KV cache□ Reserved for the OS

Qwen2.5 14B InstructQ4_K_M · 4 096 tok

VRAM
11,0 Go
10,4 Go + 0,6 Go KV
Alone
72,9 tok/s
Estimated
Concurrent
—
add a 2nd model