Meta
Llama 3.2 3B
small machines
3,2B
Parameters
1,9 Go
Q4_K_M file
≥ 3,4 Go
Recommended VRAM
131 072
Max context
Speed by graphics card
Simulate in the “Configurator” →Standard llama-bench run (prompt 512 / generation 128 tokens), model fully in VRAM. Click a card for details. Electricity cost at the default price (0,25 €/kWh).
| Card · Q4_K_M | VRAM | Generation | Prompt | VRAM used | Power | € elec. / M tok | Price |
|---|---|---|---|---|---|---|---|
| RTX 5070 Ti | 16 Go | 270 tok/s | 13 970 | 3,7 Go | 111 W | 0,03 €/M tok | — |
| RTX 4070 Ti | 12 Go | 189 tok/s | 11 590 | 3,3 Go | 126 W | 0,05 €/M tok | — |
Not measured on these cards yet — click for an estimate:
RTX 3050 Laptop GPU 4 Go should fitRTX 3050 Ti Laptop GPU 4 Go should fitRTX 3060 Laptop GPU 6 Go should fitRTX 4050 Laptop GPU 6 Go should fitRTX 3050 8 Go should fitRTX 3060 Ti 8 Go should fitRTX 3070 8 Go should fitRTX 3070 Laptop GPU 8 Go should fitRTX 3070 Ti 8 Go should fitRTX 3070 Ti Laptop GPU 8 Go should fitRTX 3080 Laptop GPU 8 Go should fitRTX 4060 8 Go should fitRTX 4060 Laptop GPU 8 Go should fitRTX 4060 Ti 8 Go should fitRTX 4070 Laptop GPU 8 Go should fitRTX 5050 8 Go should fitRTX 5050 Laptop GPU 8 Go should fitRTX 5060 8 Go should fitRTX 5060 Laptop GPU 8 Go should fitRTX 5070 Laptop GPU 8 Go should fitRTX 3080 10 Go should fitRTX 3060 12 Go should fitRTX 3080 Ti 12 Go should fitRTX 4070 12 Go should fitRTX 4070 SUPER 12 Go should fitRTX 4080 Laptop GPU 12 Go should fitRTX 5070 12 Go should fitRTX 5070 Ti Laptop GPU 12 Go should fitRTX 3080 Ti Laptop GPU 16 Go should fitRTX 4070 Ti SUPER 16 Go should fitRTX 4080 16 Go should fitRTX 4080 SUPER 16 Go should fitRTX 4090 Laptop GPU 16 Go should fitRTX 5060 Ti 16 Go should fitRTX 5080 16 Go should fitRTX 5080 Laptop GPU 16 Go should fitRTX 3090 24 Go should fitRTX 3090 Ti 24 Go should fitRTX 4090 24 Go should fitRTX 5090 Laptop GPU 24 Go should fitRTX 5090 32 Go should fit
Technical sheet
- Published
- 2024-09-18
- Architecture
- 28 layers · 8 KV heads · dimension 128
- Attention
- Standard: every layer keeps the whole context in memory.
- Context cache (f16)
- ≈ 109 MB per 1,000 tokens · 3,5 Go for 32,768 tokens
- GGUF file
- bartowski/Llama-3.2-3B-Instruct-GGUF · Llama-3.2-3B-Instruct-Q4_K_M.gguf Download (1,9 Go)
- Official repo
- meta-llama/Llama-3.2-3B-Instruct
Public Hugging Face data (official config.json, GGUF repo). The context cache only counts full-attention layers.