Gemma 3n
small machines
Pick a size: the video memory shown is the Q4_K_M file (the protocol's quantisation) plus about 1.5 GB for a 4,000-token context.
E2B
8 GB card- Parameters
- 5,4B
- Q4_K_M file
- 2,6 Go
- Recommended VRAM
- ≥ 4,1 Go
- Max context
- 32 768
Measured on 1 card · 163 tok/s (RTX 4070 Ti)
E4B
8 GB card- Parameters
- 7,8B
- Q4_K_M file
- 3,9 Go
- Recommended VRAM
- ≥ 5,4 Go
- Max context
- 32 768
Measured on 1 card · 107 tok/s (RTX 4070 Ti)
MoE (“30B-A3B”): all the memory of a 30B, but the speed of a model with 3B active parameters. “> 32 GB”: several cards, or part of the model in system memory (much slower).