Microsoft
Phi-4 14B
14,7B
Parameters
8,4 Go
Q4_K_M file
≥ 9,9 Go
Recommended VRAM
16 384
Max context
Speed by graphics card
Simulate in the “Configurator” →Standard llama-bench run (prompt 512 / generation 128 tokens), model fully in VRAM. Click a card for details. Electricity cost at the default price (0,25 €/kWh).
| Card · Q4_K_M | VRAM | Generation | Prompt | VRAM used | Power | € elec. / M tok | Price |
|---|---|---|---|---|---|---|---|
| RTX 5070 Ti | 16 Go | 84,9 tok/s | 3 994 | 10,1 Go | 182 W | 0,15 €/M tok | — |
Not measured on these cards yet — click for an estimate:
RTX 3050 Laptop GPU 4 Go too tightRTX 3050 Ti Laptop GPU 4 Go too tightRTX 3060 Laptop GPU 6 Go too tightRTX 4050 Laptop GPU 6 Go too tightRTX 3050 8 Go too tightRTX 3060 Ti 8 Go too tightRTX 3070 8 Go too tightRTX 3070 Laptop GPU 8 Go too tightRTX 3070 Ti 8 Go too tightRTX 3070 Ti Laptop GPU 8 Go too tightRTX 3080 Laptop GPU 8 Go too tightRTX 4060 8 Go too tightRTX 4060 Laptop GPU 8 Go too tightRTX 4060 Ti 8 Go too tightRTX 4070 Laptop GPU 8 Go too tightRTX 5050 8 Go too tightRTX 5050 Laptop GPU 8 Go too tightRTX 5060 8 Go too tightRTX 5060 Laptop GPU 8 Go too tightRTX 5070 Laptop GPU 8 Go too tightRTX 3080 10 Go should fitRTX 3060 12 Go should fitRTX 3080 Ti 12 Go should fitRTX 4070 12 Go should fitRTX 4070 SUPER 12 Go should fitRTX 4070 Ti 12 Go should fitRTX 4080 Laptop GPU 12 Go should fitRTX 5070 12 Go should fitRTX 5070 Ti Laptop GPU 12 Go should fitRTX 3080 Ti Laptop GPU 16 Go should fitRTX 4070 Ti SUPER 16 Go should fitRTX 4080 16 Go should fitRTX 4080 SUPER 16 Go should fitRTX 4090 Laptop GPU 16 Go should fitRTX 5060 Ti 16 Go should fitRTX 5080 16 Go should fitRTX 5080 Laptop GPU 16 Go should fitRTX 3090 24 Go should fitRTX 3090 Ti 24 Go should fitRTX 4090 24 Go should fitRTX 5090 Laptop GPU 24 Go should fitRTX 5090 32 Go should fit
Technical sheet
- Published
- 2024-12-11
- Architecture
- 40 layers · 10 KV heads · dimension 128
- Attention
- Standard: every layer keeps the whole context in memory.
- Context cache (f16)
- ≈ 195 MB per 1,000 tokens · 6,3 Go for 32,768 tokens
- GGUF file
- bartowski/phi-4-GGUF · phi-4-Q4_K_M.gguf Download (8,4 Go)
- Official repo
- microsoft/phi-4
Public Hugging Face data (official config.json, GGUF repo). The context cache only counts full-attention layers.