Card × model estimate

Llama 3.1 8B on RTX 4050 Laptop GPU Q4_K_M

too tight

On an RTX 4050 Laptop GPU (6 GB), Llama 3.1 8B in Q4_K_M needs about 6,1 Go of video memory: too tight, part of the model would have to be offloaded to system memory (much slower).

Not measured yet: the figures below are estimates, computed from the models already measured on this card. Measure my card

—

Generation (estimated)

≥ 6,1 Go

Recommended VRAM (6 GB avail.)

—

Power (ref. model)

—

Electricity / M tokens · 0,25 €/kWh

No speed estimate: the model does not fit in video memory, part of it would spill into system memory and speed would collapse (often below 5 tok/s).

What about a longer context?

ContextKV cacheEstimated VRAMFits in 6 GB?Estimated speed
4 0960,5 Go5,6 Goyes—
8 1921,0 Go6,1 Gono—
16 3842,0 Go7,1 Gono—
32 7684,0 Go9,1 Gono—
65 5368,0 Go13,1 Gono—
131 07216,0 Go21,1 Gono—

Estimated VRAM = file + KV cache (full-attention layers, f16) + ~0,5 Go of buffers, excluding memory used by the system — hence the slightly higher “recommended VRAM”, which keeps that margin. Speed ≈ estimated speed × weights read ÷ (weights read + added KV cache).

Llama 3.1 8B on other cards

RTX 3050 Laptop GPU 4 Go too tightRTX 3050 Ti Laptop GPU 4 Go too tightRTX 3060 Laptop GPU 6 Go too tightRTX 4050 Laptop GPU 6 Go too tightRTX 3050 8 Go should fitRTX 3060 Ti 8 Go should fitRTX 3070 8 Go should fitRTX 3070 Laptop GPU 8 Go should fitRTX 3070 Ti 8 Go should fitRTX 3070 Ti Laptop GPU 8 Go should fitRTX 3080 Laptop GPU 8 Go should fitRTX 4060 8 Go should fitRTX 4060 Laptop GPU 8 Go should fitRTX 4060 Ti 8 Go should fitRTX 4070 Laptop GPU 8 Go should fitRTX 5050 8 Go should fitRTX 5050 Laptop GPU 8 Go should fitRTX 5060 8 Go should fitRTX 5060 Laptop GPU 8 Go should fitRTX 5070 Laptop GPU 8 Go should fitRTX 3080 10 Go should fitRTX 3060 12 Go should fitRTX 3080 Ti 12 Go should fitRTX 4070 12 Go should fitRTX 4070 SUPER 12 Go should fitRTX 4070 Ti 12 Go should fitRTX 4080 Laptop GPU 12 Go should fitRTX 5070 12 Go should fitRTX 5070 Ti Laptop GPU 12 Go should fitRTX 3080 Ti Laptop GPU 16 Go should fitRTX 4070 Ti SUPER 16 Go should fitRTX 4080 16 Go should fitRTX 4080 SUPER 16 Go should fitRTX 4090 Laptop GPU 16 Go should fitRTX 5060 Ti 16 Go should fitRTX 5070 Ti 16 Go should fitRTX 5080 16 Go should fitRTX 5080 Laptop GPU 16 Go should fitRTX 3090 24 Go should fitRTX 3090 Ti 24 Go should fitRTX 4090 24 Go should fitRTX 5090 Laptop GPU 24 Go should fitRTX 5090 32 Go should fit
Llama 3.1 8B pageRTX 4050 Laptop GPU pageFeel the speed Download the GGUF (4,6 Go)

Estimates based on real measurements of other models on this card and on the model's public data (Meta Llama). As soon as a measurement exists, this page shows the real figures.