Alibaba
Qwen2.5 Coder 14B
code
14,8 Md
Paramètres
8,4 Go
Fichier Q4_K_M
≥ 9,9 Go
VRAM conseillée
32 768
Contexte max
Pas encore mesuré
Aucune carte n'a encore passé le protocole TokenBench sur ce modèle. Le protocole actuel couvre une sélection de modèles ; la mesure de n'importe quel modèle du catalogue depuis l'application arrive bientôt.
Mesurer ma carte →Sur quelles cartes devrait-il tenir ?
Cartes du site dont la mémoire vidéo couvre les 9,9 Go conseillés (fichier + contexte de 4 000 tokens). Estimation, pas une mesure. Cliquez sur une carte pour l'estimation détaillée.
RTX 3050 Laptop GPU 4 Go trop justeRTX 3050 Ti Laptop GPU 4 Go trop justeRTX 3060 Laptop GPU 6 Go trop justeRTX 4050 Laptop GPU 6 Go trop justeRTX 3050 8 Go trop justeRTX 3060 Ti 8 Go trop justeRTX 3070 8 Go trop justeRTX 3070 Laptop GPU 8 Go trop justeRTX 3070 Ti 8 Go trop justeRTX 3070 Ti Laptop GPU 8 Go trop justeRTX 3080 Laptop GPU 8 Go trop justeRTX 4060 8 Go trop justeRTX 4060 Laptop GPU 8 Go trop justeRTX 4060 Ti 8 Go trop justeRTX 4070 Laptop GPU 8 Go trop justeRTX 5050 8 Go trop justeRTX 5050 Laptop GPU 8 Go trop justeRTX 5060 8 Go trop justeRTX 5060 Laptop GPU 8 Go trop justeRTX 5070 Laptop GPU 8 Go trop justeRTX 3080 10 Go devrait tenirRTX 3060 12 Go devrait tenirRTX 3080 Ti 12 Go devrait tenirRTX 4070 12 Go devrait tenirRTX 4070 SUPER 12 Go devrait tenirRTX 4070 Ti 12 Go devrait tenirRTX 4080 Laptop GPU 12 Go devrait tenirRTX 5070 12 Go devrait tenirRTX 5070 Ti Laptop GPU 12 Go devrait tenirRTX 3080 Ti Laptop GPU 16 Go devrait tenirRTX 4070 Ti SUPER 16 Go devrait tenirRTX 4080 16 Go devrait tenirRTX 4080 SUPER 16 Go devrait tenirRTX 4090 Laptop GPU 16 Go devrait tenirRTX 5060 Ti 16 Go devrait tenirRTX 5070 Ti 16 Go devrait tenirRTX 5080 16 Go devrait tenirRTX 5080 Laptop GPU 16 Go devrait tenirRTX 3090 24 Go devrait tenirRTX 3090 Ti 24 Go devrait tenirRTX 4090 24 Go devrait tenirRTX 5090 Laptop GPU 24 Go devrait tenirRTX 5090 32 Go devrait tenir
Fiche technique
- Publication
- 2024-11-06
- Architecture
- 48 couches · 8 têtes KV · dimension 128
- Attention
- Classique : chaque couche garde tout le contexte en mémoire.
- Cache de contexte (f16)
- ≈ 188 Mo par 1 000 tokens · 6,0 Go pour 32 768 tokens
- Fichier GGUF
- bartowski/Qwen2.5-Coder-14B-Instruct-GGUF · Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf Télécharger (8,4 Go)
- Dépôt officiel
- Qwen/Qwen2.5-Coder-14B-Instruct
Données publiques Hugging Face (config.json officiel, dépôt GGUF). Le cache de contexte ne compte que les couches à attention complète.