Granite 4.0 H Small 32B-A9B on RTX 3080 Ti Laptop GPU Q4_K_M
too tight
On an RTX 3080 Ti Laptop GPU (16 GB), Granite 4.0 H Small 32B-A9B in Q4_K_M needs about 19,9 Go of video memory: too tight, part of the model would have to be offloaded to system memory (much slower).
Not measured yet: the figures below are estimates, computed from the models already measured on this card. Measure my card
—
Generation (estimated)
≥ 19,9 Go
Recommended VRAM (16 GB avail.)
—
Power (ref. model)
—
Electricity / M tokens · 0,25 €/kWh
No speed estimate: the model does not fit in video memory, part of it would spill into system memory and speed would collapse (often below 5 tok/s).
What about a longer context?
Context
KV cache
Estimated VRAM
Fits in 16 GB?
Estimated speed
4 096
0,1 Go
19,0 Go
no
—
8 192
0,1 Go
19,0 Go
no
—
16 384
0,3 Go
19,1 Go
no
—
32 768
0,5 Go
19,4 Go
no
—
65 536
1,0 Go
19,9 Go
no
—
131 072
2,0 Go
20,9 Go
no
—
Estimated VRAM = file + KV cache (full-attention layers, f16) + ~0,5 Go of buffers, excluding memory used by the system — hence the slightly higher “recommended VRAM”, which keeps that margin. Speed ≈ estimated speed × weights read ÷ (weights read + added KV cache).
Estimates based on real measurements of other models on this card and on the model's public data (IBM Granite). As soon as a measurement exists, this page shows the real figures.