Card × model estimate
Granite 4.0 H Tiny 7B-A1B on RTX 4070 Ti Q4_K_M
should fitOn an RTX 4070 Ti (12 GB), Granite 4.0 H Tiny 7B-A1B in Q4_K_M needs about 5,5 Go of video memory: it should fit. Estimated generation speed: ≈ 222–333 tokens/s.
Not measured yet: the figures below are estimates, computed from the models already measured on this card. Measure my card
≈ 222–333 tok/s
Generation (estimated)
≥ 5,5 Go
Recommended VRAM (12 GB avail.)
≈ 81 W
Power (ref. model)
0,02 €/M tok
Electricity / M tokens · 0,25 €/kWh
Calculation: LFM2.5 1.2B measured at 473 tok/s on this card, 0,7 Go read per token → × 0,7 Go ÷ 0,6 Go ≈ 555 tok/s (generation is memory-bound: speed follows the amount of weights read). MoE: only ~14% of the weights are read per token, but this is a theoretical upper bound; in practice llama.cpp reaches about 40 to 60% of it (expert routing, dense attention layers) → indicative range 222 to 333 tok/s.
What about a longer context?
| Context | KV cache | Estimated VRAM | Fits in 12 GB? | Estimated speed |
|---|---|---|---|---|
| 4 096 | 0,0 Go | 4,5 Go | yes | ≈ 265 tok/s |
| 8 192 | 0,1 Go | 4,6 Go | yes | ≈ 252 tok/s |
| 16 384 | 0,1 Go | 4,6 Go | yes | ≈ 230 tok/s |
| 32 768 | 0,3 Go | 4,8 Go | yes | ≈ 195 tok/s |
| 65 536 | 0,5 Go | 5,0 Go | yes | ≈ 150 tok/s |
| 131 072 | 1,0 Go | 5,5 Go | yes | ≈ 102 tok/s |
Estimated VRAM = file + KV cache (full-attention layers, f16) + ~0,5 Go of buffers, excluding memory used by the system — hence the slightly higher “recommended VRAM”, which keeps that margin. Speed ≈ estimated speed × weights read ÷ (weights read + added KV cache).
Granite 4.0 H Tiny 7B-A1B on other cards
Estimates based on real measurements of other models on this card and on the model's public data (IBM Granite). As soon as a measurement exists, this page shows the real figures.