Configurator
Build your setup, without overflowing
Pick your card, the memory to leave to the system, the model and the context size: we check that everything fits in video memory and estimate the speed.
Speeds at context 0 are measured. Anything depending on context or concurrency is an estimate: every formula is shown, nothing is written to the database.
Video memory of the RTX 4090
11,0 Go used of 23,0 Go usable (24 GB)
12,0 Go left to add a model or grow a context.
▮ model (measured)▨ KV cache□ Reserved for the OS
Qwen2.5 14B InstructQ4_K_M · 4 096 tok
- VRAM
- 11,0 Go
- 10,4 Go + 0,6 Go KV
- Alone
- 72,9 tok/s
- Estimated
- Concurrent
- —
- add a 2nd model