Google

Gemma 3

vision

Pick a size: the video memory shown is the Q4_K_M file (the protocol's quantisation) plus about 1.5 GB for a 4,000-token context.

270M

8 GB card
Parameters
0,3B
Q4_K_M file
0,2 Go
Recommended VRAM
≥ 1,7 Go
Max context
32 768

Not measured yet

1B

8 GB card
Parameters
1B
Q4_K_M file
0,8 Go
Recommended VRAM
≥ 2,3 Go
Max context
32 768

Not measured yet

4B

8 GB card
Parameters
4,3B
Q4_K_M file
2,3 Go
Recommended VRAM
≥ 3,8 Go
Max context
131 072

Not measured yet

12B

12 GB card
Parameters
12,2B
Q4_K_M file
6,8 Go
Recommended VRAM
≥ 8,3 Go
Max context
131 072

Not measured yet

27B

24 GB card
Parameters
27,4B
Q4_K_M file
15,4 Go
Recommended VRAM
≥ 16,9 Go
Max context
131 072

Not measured yet

MoE (“30B-A3B”): all the memory of a 30B, but the speed of a model with 3B active parameters. “> 32 GB”: several cards, or part of the model in system memory (much slower).