Google

Gemma 4

visionMoE

Pick a size: the video memory shown is the Q4_K_M file (the protocol's quantisation) plus about 1.5 GB for a 4,000-token context.

E2B

8 GB card
Parameters
5,1B
Q4_K_M file
3,2 Go
Recommended VRAM
≥ 4,7 Go
Max context
131 072

Not measured yet

E4B

8 GB card
Parameters
8B
Q4_K_M file
5,0 Go
Recommended VRAM
≥ 6,5 Go
Max context
131 072

Not measured yet

12B

12 GB card
Parameters
12B
Q4_K_M file
7,1 Go
Recommended VRAM
≥ 8,6 Go
Max context
262 144

Not measured yet

26B-A4B

24 GB card
Parameters
25,8B · 4B active
Q4_K_M file
15,9 Go
Recommended VRAM
≥ 17,4 Go
Max context
262 144

Not measured yet

31B

24 GB card
Parameters
31,3B
Q4_K_M file
18,3 Go
Recommended VRAM
≥ 19,8 Go
Max context
262 144

Not measured yet

MoE (“30B-A3B”): all the memory of a 30B, but the speed of a model with 3B active parameters. “> 32 GB”: several cards, or part of the model in system memory (much slower).