Models
Model catalog
The open models people run locally, organised by family, version and size. For each one: the video memory it needs, and measured speeds when it has been benchmarked.
11 families · 45 versions · 112 sizes · 6 measured
Qwen
Alibaba
The most varied family: from 0.6B for a laptop to large MoE models, with coding and reasoning variants.
9 versions · 39 sizes · 2 sizes measured
Meta Llama
Meta
The historic reference for open models: Llama 3.1 8B is the reference model of the TokenBench protocol.
4 versions · 8 sizes · 2 sizes measured
Google Gemma
Google's open models, strong at multilingual tasks; the “E” versions are designed for small machines.
4 versions · 15 sizes
DeepSeek
DeepSeek
Famous for R1 and its “distilled” versions: R1's reasoning transferred into smaller Qwen and Llama models.
5 versions · 10 sizes · 1 size measured
Mistral
Mistral AI
The French champion: efficient models, including Devstral for coding and Ministral for small setups.
7 versions · 10 sizes
Microsoft Phi
Microsoft
“Small” models trained on carefully curated data: a lot of reasoning for their size.
3 versions · 6 sizes · 1 size measured
gpt-oss
OpenAI
OpenAI's open models, MoE models shipped directly in 4-bit (MXFP4).
1 version · 2 sizes
GLM
Z.ai
Z.ai (formerly Zhipu) models, known for coding and agents; the “Flash” versions target local use.
4 versions · 5 sizes
NVIDIA Nemotron
NVIDIA
NVIDIA's models, often hybrid (Mamba + attention): a much lighter context cache.
3 versions · 5 sizes
IBM Granite
IBM
IBM's models, Apache 2.0 licensed and built for business use; Granite 4 is hybrid.
3 versions · 8 sizes
Petits modèles
Liquid AI · Hugging Face
Sub-10B models built to run anywhere, even without a dedicated graphics card.
2 versions · 4 sizes
Specs and files come from Hugging Face (official repos and public GGUFs). Speeds only come from real measurements.