Mistral AI

Mixtral

MoE

Pick a size: the video memory shown is the Q4_K_M file (the protocol's quantisation) plus about 1.5 GB for a 4,000-token context.

8x7B

32 GB card
Parameters
46,7B
Q4_K_M file
24,6 Go
Recommended VRAM
≥ 26,1 Go
Max context
32 768

Not measured yet

8x22B

> 32 GB
Parameters
140,6B
Q4_K_M file
79,7 Go
Recommended VRAM
≥ 81,2 Go
Max context
65 536

Not measured yet

MoE (“30B-A3B”): all the memory of a 30B, but the speed of a model with 3B active parameters. “> 32 GB”: several cards, or part of the model in system memory (much slower).