Microsoft
Phi-3.5 mini 3.8B
3,8B
Parameters
2,2 Go
Q4_K_M file
≥ 3,7 Go
Recommended VRAM
131 072
Max context
Not measured yet
No card has run the TokenBench protocol on this model yet. The current protocol covers a selection of models; benchmarking any catalog model from the app is coming soon.
Measure my card →Which cards should it fit on?
Cards on the site whose video memory covers the recommended 3,7 Go (file + 4,000-token context). An estimate, not a measurement. Click a card for the detailed estimate.
RTX 3050 Laptop GPU 4 Go should fitRTX 3050 Ti Laptop GPU 4 Go should fitRTX 3060 Laptop GPU 6 Go should fitRTX 4050 Laptop GPU 6 Go should fitRTX 3050 8 Go should fitRTX 3060 Ti 8 Go should fitRTX 3070 8 Go should fitRTX 3070 Laptop GPU 8 Go should fitRTX 3070 Ti 8 Go should fitRTX 3070 Ti Laptop GPU 8 Go should fitRTX 3080 Laptop GPU 8 Go should fitRTX 4060 8 Go should fitRTX 4060 Laptop GPU 8 Go should fitRTX 4060 Ti 8 Go should fitRTX 4070 Laptop GPU 8 Go should fitRTX 5050 8 Go should fitRTX 5050 Laptop GPU 8 Go should fitRTX 5060 8 Go should fitRTX 5060 Laptop GPU 8 Go should fitRTX 5070 Laptop GPU 8 Go should fitRTX 3080 10 Go should fitRTX 3060 12 Go should fitRTX 3080 Ti 12 Go should fitRTX 4070 12 Go should fitRTX 4070 SUPER 12 Go should fitRTX 4070 Ti 12 Go should fitRTX 4080 Laptop GPU 12 Go should fitRTX 5070 12 Go should fitRTX 5070 Ti Laptop GPU 12 Go should fitRTX 3080 Ti Laptop GPU 16 Go should fitRTX 4070 Ti SUPER 16 Go should fitRTX 4080 16 Go should fitRTX 4080 SUPER 16 Go should fitRTX 4090 Laptop GPU 16 Go should fitRTX 5060 Ti 16 Go should fitRTX 5070 Ti 16 Go should fitRTX 5080 16 Go should fitRTX 5080 Laptop GPU 16 Go should fitRTX 3090 24 Go should fitRTX 3090 Ti 24 Go should fitRTX 4090 24 Go should fitRTX 5090 Laptop GPU 24 Go should fitRTX 5090 32 Go should fit
Technical sheet
- Published
- 2024-08-16
- Architecture
- 32 layers · 32 KV heads · dimension 96
- Attention
- Standard: every layer keeps the whole context in memory.
- Context cache (f16)
- ≈ 375 MB per 1,000 tokens · 12,0 Go for 32,768 tokens
- GGUF file
- bartowski/Phi-3.5-mini-instruct-GGUF · Phi-3.5-mini-instruct-Q4_K_M.gguf Download (2,2 Go)
- Official repo
- microsoft/Phi-3.5-mini-instruct
Public Hugging Face data (official config.json, GGUF repo). The context cache only counts full-attention layers.