Alibaba
Qwen3 8B
raisonnement
8,2 Md
Paramètres
4,7 Go
Fichier Q4_K_M
≥ 6,2 Go
VRAM conseillée
40 960
Contexte max
Pas encore mesuré
Aucune carte n'a encore passé le protocole TokenBench sur ce modèle. Le protocole actuel couvre une sélection de modèles ; la mesure de n'importe quel modèle du catalogue depuis l'application arrive bientôt.
Mesurer ma carte →Sur quelles cartes devrait-il tenir ?
Cartes du site dont la mémoire vidéo couvre les 6,2 Go conseillés (fichier + contexte de 4 000 tokens). Estimation, pas une mesure. Cliquez sur une carte pour l'estimation détaillée.
RTX 3050 Laptop GPU 4 Go trop justeRTX 3050 Ti Laptop GPU 4 Go trop justeRTX 3060 Laptop GPU 6 Go trop justeRTX 4050 Laptop GPU 6 Go trop justeRTX 3050 8 Go devrait tenirRTX 3060 Ti 8 Go devrait tenirRTX 3070 8 Go devrait tenirRTX 3070 Laptop GPU 8 Go devrait tenirRTX 3070 Ti 8 Go devrait tenirRTX 3070 Ti Laptop GPU 8 Go devrait tenirRTX 3080 Laptop GPU 8 Go devrait tenirRTX 4060 8 Go devrait tenirRTX 4060 Laptop GPU 8 Go devrait tenirRTX 4060 Ti 8 Go devrait tenirRTX 4070 Laptop GPU 8 Go devrait tenirRTX 5050 8 Go devrait tenirRTX 5050 Laptop GPU 8 Go devrait tenirRTX 5060 8 Go devrait tenirRTX 5060 Laptop GPU 8 Go devrait tenirRTX 5070 Laptop GPU 8 Go devrait tenirRTX 3080 10 Go devrait tenirRTX 3060 12 Go devrait tenirRTX 3080 Ti 12 Go devrait tenirRTX 4070 12 Go devrait tenirRTX 4070 SUPER 12 Go devrait tenirRTX 4070 Ti 12 Go devrait tenirRTX 4080 Laptop GPU 12 Go devrait tenirRTX 5070 12 Go devrait tenirRTX 5070 Ti Laptop GPU 12 Go devrait tenirRTX 3080 Ti Laptop GPU 16 Go devrait tenirRTX 4070 Ti SUPER 16 Go devrait tenirRTX 4080 16 Go devrait tenirRTX 4080 SUPER 16 Go devrait tenirRTX 4090 Laptop GPU 16 Go devrait tenirRTX 5060 Ti 16 Go devrait tenirRTX 5070 Ti 16 Go devrait tenirRTX 5080 16 Go devrait tenirRTX 5080 Laptop GPU 16 Go devrait tenirRTX 3090 24 Go devrait tenirRTX 3090 Ti 24 Go devrait tenirRTX 4090 24 Go devrait tenirRTX 5090 Laptop GPU 24 Go devrait tenirRTX 5090 32 Go devrait tenir
Fiche technique
- Publication
- 2025-04-27
- Architecture
- 36 couches · 8 têtes KV · dimension 128
- Attention
- Classique : chaque couche garde tout le contexte en mémoire.
- Cache de contexte (f16)
- ≈ 141 Mo par 1 000 tokens · 4,5 Go pour 32 768 tokens
- Fichier GGUF
- bartowski/Qwen_Qwen3-8B-GGUF · Qwen_Qwen3-8B-Q4_K_M.gguf Télécharger (4,7 Go)
- Dépôt officiel
- Qwen/Qwen3-8B
Données publiques Hugging Face (config.json officiel, dépôt GGUF). Le cache de contexte ne compte que les couches à attention complète.